On August 7, 2026, the National Institute of Standards and Technology released the initial public draft of NIST AI 200-2, The TEVV-Athlon Framework for Evaluating AI Systems. The document addresses a gap that has become increasingly visible in organizational AI programs. Many organizations now have AI policies, governance committees, and risk registers. Far fewer can produce evidence that a specific AI system performs as intended, stays within acceptable risk boundaries, and is suitable for the environment in which it actually operates. The TEVV-Athlon Framework is NIST's proposed method for closing that gap.
TEVV stands for Test, Evaluation, Verification, and Validation. The purpose of TEVV is to generate evidence that an AI system can meet individual or organizational goals while minimizing negative impacts. The NIST AI Risk Management Framework has called for a TEVV methodology since its publication, but it did not prescribe one. NIST AI 200-2 now proposes a structured, repeatable way to design one. The draft is open for public comment through October 6, 2026.
What NIST Released
The announcement describes the TEVV-Athlon Framework as a structured approach for assessing the real-world impact and outcomes of AI systems. NIST designed it to be extensible, adaptable, and customizable, which matters because AI applications vary enormously. The framework applies to statistical machine-learning models, large language models, multimodal models, agentic systems, and other AI technologies.
The terminology deserves a moment of attention. The "TEVV-Athlon Framework" is the methodology an organization uses to design an assessment. A "TEVV-Athlon" is the resulting assessment itself. The name borrows from athletics. Just as a triathlon or decathlon evaluates an athlete across several events rather than a single test, a TEVV-Athlon evaluates an AI system through multiple events, methods, and forms of evidence rather than one benchmark or one score.
This is an initial public draft. It is not a regulation, a mandatory standard, a certification program, or a compliance requirement. It is a proposed method, and NIST is asking for input on it.
The Four-Stage Method
The framework organizes assessment design into four stages.
The first stage, Articulate & Organize, defines why the evaluation is being conducted. The organization identifies its goals, the stakeholders who need the results, the system characteristics that matter, the system's position in its lifecycle, and the resources, cost, and time the assessment will require. NIST includes seven planning questions that push evaluators to state their objectives without jargon, name who will care about the results, estimate cost and time, and confront the likely challenges before work begins. The questions function as a discipline rather than a form to fill out, because an evaluation with an unclear purpose produces evidence nobody can use.
The second stage, Define & Construct, translates high-level goals into specific measurement concepts NIST calls Metrology Blocks, or simply Blocks. A Block states precisely what the organization intends to measure and what evidence would count. Examples might include accuracy, reliability, privacy leakage, harmful output frequency, false-positive rates, false-negative rates, or robustness under abnormal inputs. The discipline here is definitional. Declaring that an AI system should be "safe," "accurate," or "trustworthy" accomplishes nothing until the organization defines how those characteristics will be observed and measured in its own context.
The third stage, Apply & Measure, determines the Events and Tools that will produce evidence for each Block. Events are the activities that generate evidence: controlled model testing, benchmark evaluations, red teaming, user testing, field testing, or testing under abnormal conditions. The Toolbox is the collection of instruments used to elicit, collect, and analyze information, from test datasets and benchmark suites to questionnaires, red-team instructions, annotation methods, and monitoring tools. The evaluation method should match the AI system, its intended use, its operating environment, and the consequences of failure.
The fourth stage, Synthesize & Interrogate, combines the evidence, examines the limitations of the measurements, and determines what the results mean for deployment, procurement, remediation, human oversight, or continued monitoring. The word "interrogate" carries weight. An organization should not simply accept a score. It should ask whether the measurement actually represents the real-world characteristic it set out to evaluate.
Where This Fits in the AI RMF
The TEVV-Athlon Framework does not replace the NIST AI Risk Management Framework. It operationalizes a specific part of it. The AI RMF organizes AI risk management into four functions: Govern, Map, Measure, and Manage. Govern and Map establish the organization's governance context, objectives, operating environment, stakeholders, and risk tolerance. Those functions feed the first stage of a TEVV-Athlon. The TEVV-Athlon itself performs the Measure function, generating evidence about system performance, risks, benefits, and impacts. The results then inform the Manage function, supporting decisions about whether to deploy, restrict, modify, monitor, remediate, replace, or reject an AI system.
The Measure function has always been the hardest part of the AI RMF to implement because it requires actual measurement rather than documentation. NIST AI 200-2 proposes the missing method.
For organizations that have adopted the AI RMF as their governance backbone, this is the practical significance of the release. NIST has separately noted that the AI RMF is being revised, but the draft does not depend on that revision.
NIST's Example: The Query-Violation Problem
The draft demonstrates the framework with what it calls the query-violation problem. A chatbot is expected to respond to a request with relevant information while avoiding a prohibited type of answer. The example measures two Blocks: helpfulness, as evidence of validity, and violation frequency, as evidence of safety. User testing and red teaming serve as the Events. The Toolbox includes user testing instructions, post-task questionnaires, red teaming instructions, and an annotation schema for labeling chats.
The lesson embedded in the example is more important than the example itself. A chatbot can be genuinely helpful and still disclose prohibited information, so measuring either characteristic alone would miss half the picture. Systems should be evaluated across multiple characteristics and operating scenarios rather than reduced to one overall score.
What This Looks Like on a Shop Floor
Consider an AI computer-vision system used to inspect machined aerospace components. This example is an application of the framework rather than one supplied by NIST, but it shows how the method translates.
The stated purpose might be to detect surface defects without rejecting an unacceptable number of conforming parts. The stakeholders include quality management, production, engineering, customers, operators, cybersecurity personnel, and executive management. The Blocks could include defect-detection accuracy, false rejections, missed defects, repeatability, performance under varied lighting and across part geometries, and behavior after a camera, model, software, or process change. The Events could include controlled tests using labeled inspection images, known edge cases, comparison against qualified human inspectors, limited production trials, abnormal lighting tests, and ongoing field monitoring.
The synthesized evidence then supports a real decision: approve the system, require human review for particular defect classes, limit its authorized use, retrain it, or reject it. That is evaluation connected to the operating reality of the business rather than to a vendor's demonstration.
The Business Implications
Several implications follow for organizations that buy, deploy, or govern AI systems, and none requires waiting for the draft to be finalized.
An AI policy is not evidence. An acceptable use policy and a governance committee define expectations, but neither demonstrates that a deployed system is accurate, reliable, secure, or suitable for its intended use. The framework's premise is that governance eventually has to produce measurement.
A vendor demonstration is not an evaluation. A polished demonstration, a generic benchmark score, or a contractual assurance may not reflect the organization's data, users, workflow, risk tolerance, or operating conditions. The framework gives buyers a vocabulary for asking what was measured, under what conditions, with what tools, and with what limitations.
Procurement will need to become more specific. Organizations evaluating AI purchases may need to define intended use, measurable acceptance criteria, testing access, system version information, change notification requirements, monitoring expectations, incident reporting, and the evidence the vendor must supply.
Evaluation must be proportional to risk. A low-impact writing assistant does not require the rigor appropriate to a system making credit decisions, detecting manufacturing defects, controlling equipment, processing sensitive information, or acting autonomously through connected tools. The framework's flexibility is designed for exactly this proportionality.
Evaluation is not a one-time activity. The draft states plainly that measurement validation is an ongoing process. Changes in the model, prompts, system instructions, retrieval data, connected tools, user population, or operating environment can invalidate prior results, so organizations should define the events that trigger reevaluation: model upgrades, vendor changes, new data sources, changes in intended use, material incidents, performance drift, or the addition of agentic capabilities.
Expertise and documentation both matter. NIST emphasizes that effective TEVV draws on technical, scientific, human-centered, and domain-specific expertise, which in a business setting means operations, quality, cybersecurity, legal, procurement, end users, and independent evaluators where appropriate. The records of an evaluation, including scope, intended use, system version, test conditions, metrics, results, limitations, decisions, and approvals, become an organizational asset that can support internal governance, customer due diligence, audits, insurance applications, and incident reviews, although they do not by themselves establish legal or regulatory compliance.
The Goodhart Warning
The draft devotes a section to Goodhart's Law, and business readers should take it seriously. Once a metric becomes the target of optimization, systems tend to be tuned to the metric rather than to the real-world outcome the metric was meant to represent. Strong benchmark performance therefore does not establish that a system is suitable for deployment, and public benchmarks carry additional problems: test data can leak into training data, and benchmarks can become contaminated or saturated in ways that inflate results. This is why NIST recommends combining multiple measurement methods and complementing controlled model testing with red teaming, user testing, and field testing under realistic conditions.
Once a metric becomes the target, the metric stops measuring what it was built to measure. Evidence of suitability comes from multiple methods under realistic operating conditions, not from a single benchmark score.
What the Framework Does Not Do
The release should not be overstated. The draft does not create a federal AI certification, a universal test that every AI system must pass, or one acceptable accuracy score for all systems. It does not replace cybersecurity, privacy, safety, contractual, sector-specific, or regulatory requirements, does not make vendor claims independently trustworthy, and does not guarantee that any AI system is risk-free. What it provides is a structured process for deciding what should be measured, how the evidence should be collected, and how that evidence should inform organizational decisions.
The Comment Period
The public-comment period closes October 6, 2026. Comments may be submitted to TEVV-Athlon@nist.gov with "NIST AI 200-2" in the subject line.
NIST states that comments may be publicly released under the Freedom of Information Act and should not contain proprietary information. NIST has specifically encouraged input from users of AI evaluation reports, including business decision-makers and procurement specialists.
The Question That Matters
The significance of TEVV-Athlon is not that NIST has produced another list of AI principles. Its significance is that NIST is proposing a repeatable way to turn governance objectives into measurable evidence and business decisions. An organization should be able to answer more than "Do we have an AI policy?" It should be able to explain what a system was expected to do, how it was tested, under what conditions, what limitations were found, who accepted the remaining risk, and what will cause the system to be reevaluated. The framework remains an initial public draft and will change before final publication, but the direction is clear.