Opens in a new tab

Evals: OpenAI’s early Framework for Evaluating LLM’s

evals openai's early framework for evaluating llm's

OpenAI Evals is an open-source framework for testing large language models and applications built with them. An eval combines a test input, an expected result or scoring rule, and a recorded outcome. It helps you answer a practical question: does this version of your AI application perform the task well enough?

The name also appears in OpenAI’s hosted evaluation tools. That distinction matters when choosing a workflow, especially because the hosted platform is being retired. This guide explains the terminology, shows a worked example, and describes how to build tests that remain useful when your tools or models change.

The current status of OpenAI Evals

Status checked on September 16, 2026. OpenAI says existing hosted evals will become read-only on October 31, 2026. The Evals dashboard and API are scheduled to close on November 30, 2026. See the current Evals documentation. This notice concerns the hosted product. The original open-source framework and the general practice of evaluating AI are separate concepts.

TermWhat it meansWhat to consider
An evalA test of model or application behaviourDefine the task and what counts as success
OpenAI Evals frameworkThe open-source project and its evaluation registryCheck its dependencies and compatibility with your model
Hosted Evals platformOpenAI’s dashboard and API for managing evaluationsPlan around the announced retirement

The original OpenAI Evals project uses evaluation definitions and datasets to run tests. Its historical workflow includes YAML configuration and JSONL data. OpenAI’s introductory cookbook is now marked as archived, so check examples before reusing installation instructions or model names.

For existing hosted workflows, OpenAI provides guidance for moving to Promptfoo. It describes manually recreating prompts, model settings, cases and scoring assertions, then validating a fresh run. Treat migration as a comparison exercise: a recreated grader may produce different scores.

Why LLM evaluations matter

A chatbot can sound convincing while routing a request to the wrong team. A document assistant can quote a relevant paragraph but attach the wrong source. A cheaper model can handle routine questions well and fail when a customer asks for a person.

Evals make these differences visible before a change reaches users. They are useful when selecting a model, changing a prompt, updating a retrieval system or adding a tool. They also give product owners something concrete to review: examples of what improved, what failed and which failures matter most.

Evaluation results depend on the test cases and scoring rules. A high score on an easy dataset offers little reassurance about a difficult production workflow. Model outputs can also vary between runs, so a single result should not carry more weight than the evidence supports.

Choosing what to measure

Start with the behaviour your application must deliver. A support router needs the right destination. A knowledge assistant needs an answer supported by the supplied documents. An agent that changes a booking needs to choose the correct action and respect the user’s approval requirements.

ApplicationUseful checksWhat a good overall score could hide
Support routingCorrect category and correct handoffRequests for a person sent back to automation
Document questionsAnswer correctness and source supportA correct answer linked to the wrong document
Structured extractionValid schema and correct field valuesValid JSON containing an incorrect amount
Tool-using agentAppropriate tool, arguments and completed taskA successful action performed without approval

Deterministic checks work well for exact labels, required fields and numeric rules. Model-based graders can assess open-ended answers against a written rubric. Human reviewers should resolve ambiguous cases and check that automated grading reflects the intended standard. OpenAI’s grader documentation explains several grading approaches, including string checks, similarity, model scoring and Python checks.

Score observable answers and actions. An explanation that sounds logical does not establish that a model’s hidden reasoning was correct.

A practical example for a support assistant

Imagine a webshop uses an LLM to route messages. The allowed labels are delivery, billing and manual_review. A request for a person must receive manual_review. For this example, messages covering both delivery and billing also go to a person.

The table contains invented cases and outputs for illustration, not measured DataNorth or model performance. Baseline A represents the current prompt. Candidate B represents a proposed update.

Customer messageExpected labelAB
Where is my parcel?deliverydeliverydelivery
The amount on my invoice is wrongbillingbillingbilling
I want to speak to a personmanual_reviewmanual_reviewdelivery
I was charged twice and my parcel is latemanual_reviewbillingmanual_review
Where is order 123?deliverybillingdelivery
Please send a VAT invoicebillingbillingbilling

Baseline A gets four of six cases right, or 66.7%. Candidate B gets five right, or 83.3%. But B fails the explicit request for a person. If respecting that request is a release requirement, B should not be released yet.

This is why an evaluation needs both an overall measure and specific failure rules. A higher average can conceal a regression that matters to your customers.

openai evals explained and how to evaluate llms

Checking the example in Python

This standalone Python example scores the candidate outputs shown above. It makes no model calls and requires no API key. Replace the outputs with responses recorded from your application when building a real test.

expected = [
    "delivery", "billing", "manual_review",
    "manual_review", "delivery", "billing",
]
candidate = [
    "delivery", "billing", "delivery",
    "manual_review", "delivery", "billing",
]
assert len(candidate) == len(expected)
correct = sum(a == b for a, b in zip(candidate, expected))
print(f"Accuracy: {correct}/{len(expected)}")
print("Human handoff passed:", candidate[2] == expected[2])

The result is Accuracy: 5/6 and Human handoff passed: False. This checks label accuracy and one handoff rule. It does not test response quality, latency or the rest of a customer conversation. Six cases are enough to explain the method, but too few to establish production reliability.

Building a useful evaluation workflow

  1. Define the decision. Specify what the test will help you decide, such as whether a new prompt can replace the current support router. Agree on critical failure rules before looking at candidate results.
  2. Assemble representative cases. Include ordinary requests, ambiguous wording, missing context and known failures. Label them with someone who understands the process. For a multilingual product, include real examples in each language instead of relying entirely on translations.
  3. Separate development from assessment. Use one set of cases to improve the prompt and reserve another for assessing the change. Record which cases were visible during development so you can recognise overfitting.
  4. Compare under consistent conditions. Keep the assessment cases and grading rules fixed while comparing versions. Record the prompt, model identifier, settings and relevant application changes. Repeat variable tasks when the decision is close.
  5. Inspect the failures. Review disagreements case by case. In the webshop example, report handoffs separately from routine routing. In a retrieval application, check the retrieved material as well as the final answer.
  6. Keep the suite useful after release. Turn newly observed failures into reviewed test cases. Run a focused suite during development and a broader suite for releases. Keep checks for older failures so improvements do not quietly reverse.

OpenAI’s evaluation best practices also emphasise task-specific tests, human calibration and ongoing evaluation. The workflow above applies those principles to a concrete release decision.

Costs and private data

Running an evaluation can involve model inference, a separate grading model, tools and human review. A useful budget starts with the number of cases, candidate versions and repeat runs. Two hundred cases tested across three variants twice require 1,200 application runs, before any separate grading calls.

Track cost and latency alongside quality. A prompt that improves accuracy but doubles the time needed to answer may be unsuitable for live support, while the same tradeoff could be acceptable for an overnight document process.

A private test repository does not mean all processing stays on your own infrastructure. Inputs sent to an external model or grader leave the local runner. Use invented examples where possible, remove unnecessary personal data, and check where prompts, outputs and evaluation logs are stored. Include both the application model and any grading service in that review.

Build an evaluation plan for your application

Choose one important workflow, write down its failure rules and compare your current application with one proposed change. If you need help defining the test cases or interpreting results, explore DataNorth’s AI consulting or discuss your evaluation project with our team.

Frequently asked questions (FAQ) about Evals

What is the difference between an eval and a benchmark

An eval is a structured evaluation. A benchmark is a standardised test or collection of tests used for comparison. Public benchmarks can help narrow a model shortlist. Your own evals determine whether a model works for your documents, users and business process.

Can I evaluate Claude or Gemini as well as OpenAI models

Yes, the evaluation method can compare outputs from different providers. The runner needs an appropriate provider integration or adapter. Confirm that it supports the model, tools and output format you intend to test, and use consistent cases and scoring rules.

How many test cases do I need

There is no universal number. Start with enough reviewed cases to cover the main behaviours, then expand where failures or uncertainty affect decisions. Report results for important categories. A hundred easy cases cannot establish reliability for an untested exception.

Do evals prevent hallucinations

They help you detect and measure unsupported answers on the cases you test. They cannot guarantee that an application will never invent information. Combine evaluation with suitable retrieval, application controls and ongoing review of real failures.

Should a new project use the hosted Evals API

Plan your evaluation workflow around tooling you can maintain beyond the hosted platform’s announced closure. Keep datasets and scoring rules portable, and review the migration guidance linked above before investing in a hosted integration.

Add DataNorth AI to your Google favorites