all the models — AI benchmark observatory
← Benchmarks

AIR-Bench

AIR-Bench 2024 is a safety benchmark grounded in risk categories derived from government regulations and company policies. It evaluates policy-grounded refusal across a broad regulatory and policy-derived harm taxonomy, using category-specific LLM-judge prompts that reward safe engagement rather than only penalizing unsafe responses.

id air-bench · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

AIR-Bench Leaderboard · all the models