all the models — AI benchmark observatory

News

Live AI headlines merged from lab blogs and TechCrunch — refreshed about every 15 minutes.

10 stories from Anthropic

  1. AnthropicSep 22, 2026, 4:00 PM UTC
    Claude discovers a novel enzyme system with CRISPR-like repeats

    We’re introducing a new life sciences research group and laboratory at Anthropic. Our focus is on fundamental biology research using Claude: exploring datasets of DNA to identify uncharacterized protein families, generating hypotheses at scale, and testing them through experiment

  2. AnthropicSep 17, 2026, 4:00 PM UTC
    Partnering with Accenture on embedded evaluation

    We're partnering with Accenture on independent evaluation of frontier AI. This is an important step toward the commitment, made in our CEO’s essay “We Must Pace the Frontier,” to embed evaluators within Anthropic. The partnership will be led by Faculty, Accenture’s specialist AI

  3. AnthropicSep 16, 2026, 4:00 PM UTC
    Introducing the Life Sciences Verification Program

    Today, we are introducing the Life Sciences Verification Program (LSVP), which gives life science professionals access to our Mythos, Opus, and Sonnet models with a refined set of safeguards more permissive for biology-related work. We have already onboarded dozens of organizatio

  4. AnthropicAug 31, 2026, 4:00 PM UTC
    Developing Enterprise Frontier Safeguards with our customers

    Today we’re announcing Enterprise Frontier Safeguards (EFS), a solution that combines the privacy of zero data retention (ZDR) with state-of-the-art safeguards for detecting misuse. EFS works by storing data in cloud infrastructure controlled by the customer, not Anthropic. EFS w

  5. AnthropicAug 30, 2026, 4:00 PM UTC
    Improving our alignment and security efforts

    On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation en

  6. AnthropicAug 26, 2026, 4:00 PM UTC
    Expanding our support for scientists

    As Claude becomes increasingly capable at scientific research, we are focusing on building products and programs to support the research community. In June, we launched Claude Science , a product that integrates the tools that researchers most commonly use, produces auditable art

  7. AnthropicAug 26, 2026, 4:00 PM UTC
    Previewing the Model Hardware Standard

    We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first group of scientific research labs and advanced manufacturers. MHS enables AI agents to operate multiple lab and manufacturing

  8. AnthropicAug 24, 2026, 4:00 PM UTC
    Funding better evaluations of AI’s impact on wellbeing

    We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how ou

  9. AnthropicAug 11, 2026, 4:00 PM UTC
    How Claude’s text watermark works

    Future Claude models will generate text that contains a watermark. This is a way of determining the likelihood that Claude was involved in writing the text, and we, along with several other major AI providers, are implementing this change to comply with the EU AI Act. In this art

  10. AnthropicAug 6, 2026, 4:00 PM UTC
    Improving Fable 5's biology safeguards

    We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer “fallbacks”—where the system switches to a less capable model after they make a biology-related query. In our testing, thi