Use a prototype when the team needs to learn whether an approach can work at all under bounded conditions. Use a pilot when the team needs to learn whether a more complete capability can create value with representative users, workflow, integration, and oversight. The next decision should determine which evidence comes first.
These are working definitions, not universal acquisition terms. A program should align its terminology and artifacts with the applicable authority and pathway. The important distinction is the type of uncertainty each activity is designed to reduce.
What is an AI prototype supposed to prove?
An AI prototype should answer one or a small number of consequential questions. It might test whether a model can perform a bounded task, whether a local inference stack can meet latency requirements, whether an agent can use an approved tool under restricted permissions, or whether representative data supports the proposed approach.
The prototype does not need every production control, integration, or interface. It does need a credible test boundary, recorded configuration, representative cases, explicit measures, failure conditions, and a decision owner. Otherwise, the result may be impressive without reducing the uncertainty that justified the work.
GAO’s review of defense prototyping found that programs used prototypes to reduce technical risk, investigate integration challenges, validate designs, mature technology, and refine performance requirements. In that review, prototyping made business cases more realistic by providing evidence about maturity, feasibility, cost, and achievable performance before larger commitments. SeeGAO-17-309 on defense prototyping.
What is an AI pilot supposed to prove?
An AI pilot should test a more complete capability in a representative workflow. The open questions usually concern mission utility, human factors, integration, oversight, operating burden, support, and behavior over a broader range of realistic conditions. The pilot asks whether users and the organization can work with the capability, not only whether the core technology functions.
A pilot therefore needs more organizational readiness. Representative users must participate. Data and integration paths must be approved for the test. Human review, escalation, logs, support, and incident handling need enough definition to make the observations meaningful. The pilot should still be controlled and reversible. It is not a substitute for the approvals required for operational deployment.
| Dimension | Prototype | Pilot |
|---|---|---|
| Primary question | Can this bounded approach answer the central technical or design assumption? | Can this capability create useful outcomes in a representative workflow? |
| Scope | Smallest credible experiment | Controlled but more complete user and system experience |
| Users | Subject matter input and limited test participation | Representative users performing realistic tasks |
| Environment | Isolated lab, simulation, synthetic or approved test data | Representative environment with approved integrations and controls |
| Evidence | Feasibility, limits, performance, failure behavior, integration risks | Mission utility, human factors, workflow fit, support burden, broader reliability |
| Next decision | Advance, revise, test another assumption, or stop | Prepare for a defined transition, conduct further evaluation, revise, or stop |
Which evidence should come first?
Test the uncertainty that could most change the decision and can be investigated safely at the smallest credible scale. If the team does not know whether the core task is technically feasible, do not start with a pilot. If feasibility is already credible but workflow and human oversight remain unknown, another model-only prototype may add little value.
The Department’s Prototyping and Experimentation organization describes a body of evidence that includes operational utility, technical feasibility, integration risk, user feedback, and transition pathways. Its public overview emphasizes structured prototyping, field experimentation, mission-relevant environments, and evidence-based transition decisions. See the officialPrototyping and Experimentation overview.
This body-of-evidence view is more useful than asking whether a demo “worked.” Each stage should add evidence that the prior stage could not produce. The record should also preserve failure, user intervention, configuration, assumptions, and limits so the next decision is not built on a simplified success story.
How do you choose between a prototype and a pilot?
Choose by examining the central uncertainty, the maturity of the technical approach, the readiness of the workflow, and the scope of the next decision. If several foundational conditions are still unknown, separate them instead of asking one large pilot to answer everything.
- Choose a prototype when feasibility, model behavior, data suitability, architecture, or a high-risk integration is still uncertain.
- Choose a pilot when the approach is credible and the decision requires evidence about users, workflow, oversight, integration, and operating value.
- Choose discovery first when the mission decision, user, owner, or success condition is unclear.
- Choose evaluation first when several technologies could meet a defined need but the comparison criteria and evidence are incomplete.
The Mission AI Readiness Worksheetcan help reveal whether the opportunity has enough mission, data, technical, governance, evaluation, and adoption context for a bounded test. It deliberately distinguishes readiness to test from readiness to deploy.
What belongs in the test plan before work begins?
A short test plan should make the decision logic visible. It does not need to predict every engineering detail. It needs to prevent the scope, measures, and interpretation from shifting after the team sees the result.
- Decision question: What choice will the evidence support?
- Central assumption: What must be true for the concept to remain credible?
- Boundary: Which users, cases, data, tools, environment, and timebox are included?
- Measures: Which observable outcomes matter, and how will they be calculated?
- Acceptance and failure: What supports advancing, and what requires stopping or redesign?
- Oversight: Who may intervene, review, approve, or halt the test?
- Evidence package: Which configuration, results, limits, and observations will be preserved?
CDAO’s public AI testing frameworks separate model evaluation, human systems integration, system integration, and operational testing. That structure helps a team identify which layer the current experiment can support and which conclusions must wait. Review theCDAO test and evaluation framework summarybefore treating one test type as evidence for another.
What are the most common failure patterns?
The first failure is building before defining the decision. The team produces a capability, then searches for measures that make it look useful. The second is confusing a scripted demonstration with repeated performance. The third is expanding users, data, integrations, and permissions before the core assumption has survived a bounded test.
Another failure is the permanent prototype. It accumulates interfaces, exceptions, and expectations without a transition owner, support model, evidence threshold, or stop decision. A prototype should end with a decision. If it becomes the operational solution by inertia, the program inherits risks that were intentionally left outside the experimental scope.
Finally, teams often preserve positive output but lose the test context. Developmental test and evaluation is intended to generate information about capabilities and limitations. The officialDevelopmental Test and Evaluation overviewdescribes this as deliberate exercise and evaluation across the lifecycle. Configuration, conditions, failures, and limits are part of the evidence.
When should the program transition to a broader effort?
Transition when the evidence answers the current question, the next uncertainty requires a broader setting, and a responsible owner accepts the expanded scope. Before a pilot, confirm representative users, approved data and integrations, oversight, monitoring, support, measures, stop conditions, and a decision date.
A transition is not always forward. The right result may be a revised concept, a different technology, more discovery, or a stop decision. Evidence-first prototyping creates value by reducing uncertainty, even when it prevents a larger investment.
Frequently asked questions
How long should an AI prototype take?
It should be timeboxed to the smallest period that can answer the central question. The right duration depends on data, integration, environment, and test complexity. If scope keeps growing, split the assumptions and choose the first decision gate.
Does a pilot use operational data?
Not automatically. Data use must follow applicable handling, privacy, security, legal, and authorizing requirements. A pilot can use approved representative data or other bounded arrangements, but the evidence must state what those conditions allow the team to conclude.
Can a successful prototype move directly to deployment?
A prototype alone rarely establishes workflow, human factors, system integration, monitoring, support, security, governance, and lifecycle suitability. The applicable program and authorizing authorities determine the required pathway and evidence before operational use.