Understanding the U.S. Federal Government’s AI Spending
Photo: Bettmann / Contributor via Getty Images
The U.S. government is buying more artificial intelligence now than ever before. Between fiscal year 2019 and FY 2025, federal agencies awarded 2,255 identified AI-related awards, with $4.1 billion in total obligations. The number of AI-related contracts more than tripled over this period, and annual obligations also increased significantly, reaching $1.162 billion in FY 2025. This increasing demand is dominated by the Department of Defense (DOD), which accounts for 83 percent of obligations. The CSIS Futures Lab followed the federal money trail to understand how the U.S. government buys artificial intelligence, drawing on two sources: solicitation notices posted on SAM.gov and awarded federal contracts recorded in USAspending.
Q1: Is federal demand for AI increasing?
A1: Yes. Federal demand for AI has increased substantially since FY 2019 and continues to grow. Figure 1 shows that the number of awarded contracts mentioning artificial intelligence more than tripled between FY 2019 and FY 2025. SAM.gov solicitation notices also increased, with FY 2025 reaching 71 AI-related notices.
In terms of total obligations, annual federal AI obligations rose sharply, reaching $1.162 billion in FY 2025. Among these contracts, generative AI is also becoming increasingly visible. Following ChatGPT’s launch in November 2022, contracts tagged as generative AI climbed to 56 in FY 2024 and 118 in FY 2025 (see Figure 2).
Future demand and federal spending on AI are likely to continue to increase. Existing announcements show that the government will increase spending to modernize AI infrastructure. Data centers are a useful indicator, given the scale of investment flowing into this infrastructure. Federal AI demand and data center growth were correlated until the LLM revolution; however, since ChatGPT’s launch, data center growth has far outpaced federal AI contracting. This suggests that AI infrastructure investments are now mostly driven by commercial demand, and that federal spending is lagging behind the private sector on AI investments.
Q2: Who is getting AI contracts?
A2: On the supply side, many suppliers participated while a substantial share of funding went to a few recipients. Across 2,255 AI-related contracts, 1,250 different vendors won awards (see Figure 4). Of those vendors, 66.2 percent won only one contract. At the same time, the three largest recipients accounted for 30.4 percent of net obligations, reflecting both the top-heavy and long-tailed nature of the market. The median award value is about $150,000 per year. Small businesses win most AI contracts by count (70 percent), but they hold a considerably smaller share of the total obligations (45 percent), meaning that they take many lower-value awards (see Figure 5).
This data suggests that the federal AI market remains relatively fragmented, while coalescing around a few heavyweights near the top. Most agencies are using many small and mid-sized awards. It is likely that the government is testing different tools, vendors, and use cases before committing to larger programs.
Large IT and consulting firms remain important within federal AI acquisitions. Traditional defense primes are less prominent than expected, accounting for just 2.8 percent of contracts and 5 percent of obligations. Some commercial AI firms have a larger dollar impact despite a small number of contracts: Palantir and Scale AI together account for $597 million, about 15 percent of obligations from less than 1 percent of contracts.
Q3: Which departments and agencies drive AI acquisition?
A3: The DOD is the central actor in federal AI acquisition, accounting for roughly 64 percent of AI contracts and 83 percent of obligations between FY 2019 and FY 2025 (see Figure 6). Civilian agencies such as the Department of Health and Human Services (HHS), the Department of Homeland Security, and the General Services Administration each account for much smaller shares. On the civilian side, HHS is the largest buyer by contract count (296 contracts, including 131 with the National Institutes of Health), followed by NASA (219 contracts), the Department of Transportation (136), the Department of Commerce (129), and Homeland Security (122).
Within the DOD (Figure 7), the Air Force is the largest AI buyer in the awarded contract data, with 1,182 contracts and $872 million in obligations. The Defense Advanced Research Projects Agency also appears as a significant buyer, with 187 contracts worth $213 million. However, both are much less visible in the solicitation data used for this analysis. Furthermore, public data likely understates defense-related acquisition, of AI as some vehicles are less visible in open solicitation records. Moreover, AI spending within the intelligence community is likely not captured in the data presented here. This data should be considered as an indicative of a larger trend, as real spending numbers may be considerably higher.
Q4: How is AI being governed?
A4: AI needs to be governed carefully, as it could have significant impact on national security. Models need careful evaluation before integration into government workflows; for example, many AI models show risky tendencies when performing key military tasks. Any acquisition process should detail key AI governance frameworks such as monitoring, data rights, benchmarking, model portability, and incident reporting.
However, according to the full solicitation language of AI-related notices for governance requirements, AI governance is uneven and limited. Overall, over half of AI solicitation notices in the sample contain no clear governance language. As shown in Figure 8, most AI-related contracts fall into two broad categories: “machine learning” and “general AI.” The “generative AI” category is growing quickly but remains a smaller share of total federal AI spending. Because many federal contract descriptions are short, the generative AI counts should be read as a lower-bound estimate.
Among the notices that do include governance language, most focus on performance measurement and process requirements, as can be seen in Figure 9. Benchmarking appears in over one-third of notices. These provisions can help agencies measure whether an AI system performs as promised, but they do not necessarily create strong accountability if a system fails or produces harmful outcomes. Hard accountability clauses such as red teaming are much rarer. Audit access, red teaming, and incident reporting each appear in fewer than 10 percent of notices. The exception to the scarcity of hard accountability rules is a relatively high incidence of data rights provisions, which spell out who owns and can use the data and models involved.
When agencies do include governance language, benchmarking and evaluation provisions are among the most frequently documented forms of governance language, appearing in 28 notices (37 percent). These provisions specify performance tests, evaluation metrics, or acceptance criteria. Explicit monitoring provisions, requiring repeated tracking or review of AI performance during development or operation, were confirmed in four notices (5 percent). With respect to co-occurrence, benchmarking and data rights provisions are the most frequently documented combination in the reviewed sample, appearing together in 23 of 75 notices (31 percent). These provisions address the evaluation of performance and the allocation of rights in specified data, software, or other deliverables. Benchmarking and model portability provisions appear together in 17 notices (23 percent), while benchmarking and monitoring were confirmed together in three (4 percent). Incident reporting requirements are the rarest, appearing in only three notices. In short, governance language in federal AI acquisition is inconsistent: A few agencies write comprehensive requirements, many include only basic performance measures, and over half include nothing at all.
Overall, there has been a significant increase in the U.S. government’s AI acquisition, mostly driven by defense-related contracts. U.S. AI strategy relies on the country’s continued AI leadership; therefore, the government should continue spend on AI at the federal level. As federal AI spending grows, agencies should use procurement to translate that investment into quantifiable public value creation, complete with clear expectations for performance and accountability.
Yasir Atalan is deputy director and data fellow in the Futures Lab at the Center for Strategic and International Studies in Washington, D.C. Erik Tiersten-Nyman is a research associate in the Futures Lab at CSIS. Benjamin Jensen is director of the Futures Lab and a senior fellow in the Defense and Security Department at CSIS.
Appendix
We analyze publicly reported USAspending prime contracts and orders with selected activity in FY 2019–FY 2025, identifying AI-related purchases through descriptions containing AI terminology and documentary review of named programs. There are 12 supported program awards/orders, including Maven Smart System, Smart Sensor, CBC2, and Advana Edge, that add $736.35 million without overlap with the keyword sample. The combined sample contains 2,255 distinct awards/orders, represented by 3,602 award-year observations and 5,874 transactions totaling $4.104 billion in nominal net obligations. These figures cover federal contracts only. Grants and Other Transaction Agreements are excluded, so the totals are best read as a lower bound on federal AI activity. Annual counts include previously awarded contracts active that year; the headline counts each award/order once across the study period. Obligations retain negative adjustments and represent commitments, not payments or potential contract ceilings. Mixed-purpose purchases contribute their full selected award-year obligations. Grants, other transaction agreements, subcontracts, parent vehicles, and facility construction are excluded. Dollar shares follow transaction-level funding agencies and business size determinations; annual count attributes follow the latest transaction. Figure 6 counts use awarding agencies, while Figure 7 uses funding organizations. Business size is distinct from set-aside status. A separate vendor name screen identified approximately $3.995 billion in candidate records, but these were excluded because supplier identity alone does not establish AI content and the pool contains parent vehicles and overlap. This amount is not a validated addition. Omissions, possible false inclusions, and mixed-purpose costs mean our total is neither comprehensive nor a formal lower bound on federal AI spending.
Governance analysis covers the 75 Solicitation or Combined Synopsis/Solicitation records in the original 278-notice collection, selected by notice type rather than clause findings. This is a nonrandom review frame, separate from the financial sample; related notices remain separate observations. Available material comprises attachment text for 46 records, notice text alone for 23, and no resolved usable text for 6. Scope review retained 62 records, excluded 6, and left 7 unresolved; all 75 remain in the reported denominator. Keyword searches and supplementary terms identified passages for contextual review across eight categories: benchmarking, data rights, portability, audit access, human oversight, red teaming, monitoring, and incident reporting. Each notice counts at most once per category. Definitions include data/interface portability, bias and robustness assessment, and some development-stage safeguards; incorporated or proposed provisions do not establish final negotiated requirements or implementation. Percentages are minimum documented shares of this review frame, not estimates of government-wide prevalence. At least one provision was confirmed in 34 notices; the remaining 41 have none confirmed, which does not establish absence. Co-occurrence measures provisions documented in the same notice, not deliberate bundling. Incomplete source packages, text-extraction limitations, and the absence of independent second-human coding constrain interpretation.