Answer capsule
NIST's August 2026 initial public draft introduces TEVV-Athlon as a four-stage way to build customized assessments around organizational test, evaluation, verification, and validation objectives. The framework is meant to accommodate many technologies and contexts, and NIST is seeking input through October 6. That flexibility does not tell a CIO what evidence is sufficient for a production decision. Before funding an evaluation, the CIO should approve the exact decision, claims, operating context, measures, and limitations the resulting evidence must support.
What the source establishes
- NIST announced the initial public draft of NIST AI 200-2, the TEVV-Athlon Framework, in August 2026 and is accepting input through October 6, 2026.
- The page describes a four-stage method for developing customized AI assessments based on organizational TEVV objectives.
- NIST says the approach is intended to be extensible across statistical machine learning, large language, multimodal, agentic, and other AI systems and contexts.
- The initial draft does not choose a buyer's decision, claims, operating context, measures, acceptance thresholds, production authority, or retest conditions.
Name the decision the assessment must inform
State whether the evidence will support research, vendor comparison, design selection, limited pilot, production release, authority expansion, continued operation, incident response, or retirement. Name the accountable decision owner, affected users and stakeholders, system and version, operating environment, consequence of error, claims being tested, and claims explicitly outside scope. The CIO should reject a generic objective such as evaluate the model because the same score can mean different things for a read-only assistant, an automated action, or a high-consequence decision. Customization begins with a decision boundary, not a menu of available benchmarks.
Turn each claim into observable evidence
For each capability, performance, safety, reliability, security, fairness, privacy, cost, usability, or outcome claim, define the relevant event or scenario, system behavior, measurement concept, data source, comparison, threshold, uncertainty, and reviewer. Include expected work alongside missing context, conflicting instructions, distribution shifts, dependency failures, permission changes, adversarial inputs, human overrides, and recovery. Preserve the exact system, configuration, test material, tool, and result. An evaluation should show what was observed and under which conditions; a composite score should not erase a failed critical scenario or imply performance outside the assessed context.
Protect the meaning of the result
Separate the team that configures the system, the team that designs and runs the assessment, the experts who interpret domain consequences, and the executive who accepts the decision. Record reused test data, optimization against known cases, missing populations, assessor access, manual interventions, excluded failures, and conflicts. Give business and affected-user reviewers enough detail to challenge whether the test represents real work. A vendor result, laboratory result, or public benchmark can inform the evidence plan, but it cannot automatically be transferred to the buyer's assembled system, data, integrations, users, and operating controls.
Bind acceptance to limits and retest triggers
The acceptance record should state which claims passed, failed, or remain unknown; which risks and exceptions were accepted by whom; the permitted users, data, actions, volume, and duration; and the monitoring, fallback, pause, and retirement conditions. Name changes that invalidate or narrow the result, including a model, prompt, data, tool, permission, interface, provider, policy, user population, or purpose change. Reopen the evidence plan when the draft framework changes if that change matters to the decision. Completion of a TEVV exercise is not production approval, and an initial NIST draft does not establish universal sufficiency or transfer accountability from the CIO.
Turn this source into a reviewable decision
For AI for CIOs, use this briefing as a dated decision record rather than a substitute for the source. Preserve The TEVV-Athlon Framework for Evaluating AI Systems, the exact URL, the September 4, 2026 review date, the supported facts above, the editorial interpretation, the limitations, and any buyer-specific evidence. Link that record to the decisions most directly affected: AI portfolio economics; Operations and incident intelligence; Software delivery and modernization; Enterprise AI platform architecture. State whether the source changes the scope, evidence requirement, control, sequence, or only the language used to describe the decision.
Before action, name the accountable owner, affected population and workflow, exact offering or configuration, source data and rights, human decision point, exception and appeal path, complete cost, expected benefit, failure and stop conditions, retained evidence, and next review date. Keep official facts, provider statements, buyer observations, representative tests, measured outcomes, editorial inferences, and unknowns visibly separate. Reopen the record when the source, offer, model, integration, data, policy, population, responsible person, or measured result changes.
Limitations and unknowns
NIST is the primary government source. The page presents NIST AI 200-2 as an initial public draft, describes a four-stage customizable TEVV-Athlon approach, and seeks comments through October 6, 2026. It does not provide a final standard, universal evaluation plan, buyer-specific objective, test population, measure, acceptance threshold, assessor qualification, production decision, certification, or proof of capability, security, safety, reliability, fairness, usability, cost, impact, or outcome. The current draft and any later revision, exact system and operating-context record, stakeholder and claim analysis, representative test design and data, versioned results and limitations, independent review where warranted, production observation, change evidence, and qualified architecture, engineering, evaluation, domain, security, privacy, accessibility, procurement, business-owner, regulatory, and legal review control.
Decision test
Ask whether the source changes the decision itself, the evidence required, the implementation sequence, or only the language used to describe an existing capability. Record which claims are directly supported, which are provider statements, which require an independent test, and which remain unknown. A source-linked review should make uncertainty easier to see, not bury it inside a blended score.
Questions to take into review
- What is the unit of useful work?
- How does cost change with context, retrieval, tool calls, retries, and review?
- Which telemetry is missing or sampled?
- Can the model change production or only advise?
- Which repositories and dependencies are exposed?
- What checks gate generated changes?
- Which services are common and which remain workload-specific?
- How can a team change a model without rewriting the application?
The publication supports research and executive decision preparation. It does not provide legal, financial, accounting, employment, clinical, cybersecurity, investment, procurement, or implementation advice.