← Benchmarks
ResearchClawBench
ResearchClawBench evaluates research agents on realistic, tool-using research tasks that require code execution and filesystem workspace interaction.
id researchclawbench · max 1 · 1 models reported
| # | Model | Score |
|---|
No scores for this benchmark yet.
