all the models — AI benchmark observatory
← Benchmarks

VIBE-V2

VIBE-V2 is an internal benchmark covering pure front-end and full-stack Web, Android, and iOS projects with build-from-scratch tasks. It uses an Agent-as-a-Verifier paradigm to automatically verify program interaction logic and visual output, scoring models through a unified pipeline that includes a requirement set, containerized deployment, and a dynamic interaction environment.

id vibe-v2 · max 1 · 1 models reported

#ModelScore

No scores for this benchmark yet.

VIBE-V2 Leaderboard · all the models