On September 22, a senior Pentagon research and engineering official offered a number to show how far the Maven Smart System has come. “In January of this year, we had about 50,000 people using Maven,” James Mazol told a DefenseScoop audience, and after the start of Operation Epic Fury the figure had passed 100,000. The number is real evidence of something. It shows demand, and it shows that commanders found the tool worth logging into during a war. It does not say how often the system’s detections were correct, under what conditions they were wrong, or whether the version in use last month behaves like the one in use today.
That distinction matters because Maven is no longer an experiment. The department raised the ceiling on Palantir’s contract from $480 million to nearly $1.3 billion in May 2025, and a March 9, 2026 memo from Deputy Secretary Steve Feinberg directed its transition to a program of record by the end of the fiscal year. Yet when analysts at the Center for Strategic and International Studies assembled the public case for the system in June, the headline performance claims traced back to newspaper reporting on the Iran campaign, an unnamed official quoted in a book, and a Georgetown assessment of 18th Airborne Corps exercises from 2023 and 2024. None of those is a test report.
A program office buying an AI-enabled system is buying a stream of evidence that has to be renewed every time the software changes. The practical discipline is to know which kind of evidence stands behind each claim, and to write the renewal into the contract before award. Recent policy has made that harder to do and more necessary.
A demo, a benchmark and a trial answer different questions
A demonstration shows that a system can perform a task once, on inputs the vendor selected, usually with the vendor’s engineers nearby. A laboratory benchmark shows how a model scores on a fixed dataset, which is useful for comparing models and nearly silent on how the fielded system performs when the inputs stop resembling that dataset. An operational trial puts the whole system, including its operators, in representative conditions and counts outcomes. Usage statistics are a fourth category, and they measure adoption.
The department’s own testers are candid about how weakly these transfer. The developmental test guidebook for AI-enabled systems, published by the research and engineering office in February 2025, lists the sensitivity of models to small input changes and their dependence on training data, then concludes that these factors “undermine the ability of test teams, evaluators, and executives to generalize results from particular tests.” It also quotes a Chief Digital and AI Office framework conceding that “it is not feasible for OT&E to completely cover the operational envelope.”
Maven supplies the standard illustration. Bloomberg’s 2024 reporting, as summarized by The Batch, put the system’s success rate at identifying objects at about 60 percent against 84 percent for human analysts at the 18th Airborne Corps, and noted that its training data emphasized deserts, with performance falling in other environments. CSIS says performance has improved considerably since the exercise-era assessments. A buyer should ask on what test, against which imagery, and with which software build.
What the solicitation should pin down
The first question is the intended task and the conditions it must be performed in, stated narrowly enough that a test could fail. “Detects vehicles” is a capability description. A requirement names the sensor, the terrain, the weather, the clutter and the acceptable false-alarm rate. The Government Accountability Office’s April 2026 review of federal AI acquisitions, which examined four Defense Department efforts including Maven, found that Maven officials learned their early contracts “lacked AI-related requirements needed to hold vendors accountable” and tightened them in later awards.
The second is the test population. Any accuracy figure is a statement about the data it was computed on. The program office should know who assembled the test set, whether the vendor had access to it during training, and how closely it resembles the theater the system is headed for. A government-held test set that the vendor has never seen is worth more than any number in a proposal.
Human intervention is the third. Reported results often describe a team of machine and operator, and the operator’s corrections vanish into the aggregate. Buyers should ask how many outputs were overridden, how long each review took, and what the system does when no reviewer is available. DoD Directive 3000.09, as revised in January 2023, requires that autonomous and semi-autonomous weapons allow “appropriate levels of human judgment” over the use of force, demands rigorous verification and realistic operational testing, and sets senior-level approvals before formal development and again before fielding. That directive is itself in motion: National Security Presidential Memorandum 11, signed June 5, 2026, ordered an update within 90 days and annual reviews afterward. Program offices should confirm which text governs their system rather than assume.
Failure cases come last and are asked for least. A vendor that cannot describe where its system breaks has either not looked or would prefer not to say.
Every update reopens the evidence
Since March 2025 the Software Acquisition Pathway has been the preferred route for software development in business and weapon programs, with commercial solutions openings and other transactions as the default instruments. The January 2026 AI strategy went further, directing a vendor cadence that puts the latest commercial models into use within 30 days of public release and creating a monthly board empowered to waive nonstatutory requirements.
Speed of that kind is defensible only if validation keeps pace. The test guidebook warns that fielded performance can diverge from tested performance and that drift “can also be caused by frequent AIEC software updates”; it recommends performance thresholds after fielding that trigger intervention. An academic analysis of the software pathway posted in June 2026 found that AI-specific controls for data provenance, lifecycle management and human oversight sit in supplementary documents instead of the artifacts program offices actually work from. NSPM-11 gives the department 120 days to propose standard test, evaluation, verification and validation methods for national security AI, which is welcome and not yet a method.
Independent capacity to check has meanwhile shrunk. A May 2025 memo cut the operational test director’s office from 94 personnel to 46 and ended its contractor support. GAO reported on June 30, 2026 that the office’s action officers were covering more programs, some outside their expertise, across 173 systems on the oversight list. The burden of asking hard questions therefore falls more heavily on the program office than it did two years ago.
Accountability and data rights are contract terms
An AI-enabled system usually has several authors: a model developer, an integrator, a cloud provider and a government program office. Organization charts have also moved. The Chief Digital and AI Office was placed under the research and engineering undersecretary in August 2025, and the Feinberg memo shifts Maven’s administration from the National Geospatial-Intelligence Agency to a new program office there while its contracts move to an Army enterprise vehicle. The November 2025 acquisition overhaul promised a more competitive vendor base and less lock-in. Each change is reasonable on its own. Together they make it easy for nobody in particular to own the question of whether last week’s model update was revalidated.
The model layer can also change for reasons unrelated to performance. On February 27, 2026 the department designated Anthropic a supply chain risk after the company declined the “any lawful use” contract language the AI strategy requires; according to Lawfare’s account, reporting indicated its model had been deployed on classified networks through Maven. Whatever one thinks of the dispute, it shows that a component inside a fielded system can be withdrawn or replaced on short notice, and the evidence gathered on the old component does not carry over.
Data rights determine whether the government can respond. GAO found the Maven program had trouble deciding before award what intellectual property and data rights it needed “to promote future competition,” and officials told auditors the government should prioritize owning data over owning algorithms. GAO’s June 2023 recommendation for department-wide AI acquisition guidance is still listed as open and only partially addressed. The NIST AI Risk Management Framework offers a usable vocabulary in the meantime, though it is voluntary and under revision.
Three changes would cost little. Solicitations should require vendors to label every performance claim by the kind of evidence behind it and the software version tested. Contracts should give the government the test sets, labeled operational data and logs needed to rerun an evaluation without the vendor, and should name the conditions, including any model substitution, that trigger one. And each program should designate a single official who signs for revalidation after an update, with the results filed in the lessons-learned repository that GAO has asked the department to feed, a recommendation the department accepted.
This analysis draws on the public sources linked in the text. Send corrections to [email protected].


