Introducing SWE-bench Verified
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
Topics: Models
Entities: Models
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
Topics: Models
Entities: Models
Only models with a post-mitigation score of "medium" or below can be deployed.Only models with a post-mitigation score of "high" or below can be developed further.
Topics: Models
Entities: Models
Rakuten Pairs Data with AI to Unlock Customer Insights and Value
We are introducing Structured Outputs in the API—model outputs now reliably adhere to developer-supplied JSON Schemas.
We’re testing SearchGPT, a temporary prototype of new search features that give you fast and timely answers with clear and relevant sources.
We've developed and applied a new method leveraging Rule-Based Rewards (RBRs) that aligns models to behave safely without extensive human data collection.
Topics: Models
Entities: Models
OpenAI is committed to making intelligence as broadly accessible as possible. Today, we're announcing GPT‑4o mini, our most cost-efficient small model. We expect GPT‑4o mini will significantly expand the range of applications built with AI by making...
Topics: Models
Compliance API integrations, SCIM, and GPT controls to support compliance programs, data security, and user access at scale
Topics: Models
Entities: ChatGPTModelsChatGPT Enterprise
Discover how prover-verifier games improve the legibility of language model outputs, making AI solutions clearer, easier to verify, and more trustworthy for both humans and machines.
OpenAI and Los Alamos National Laboratory are working to develop safety evaluations to assess and measure biological capabilities and risks associated with frontier models.
Topics: Models
CriticGPT, a model based on GPT-4, writes critiques of ChatGPT responses to help human trainers spot mistakes during RLHF
Topics: Models