Ninety percent. That is the approximate share of paid AI spend in the House that flows into OpenAI’s coffers. While the press likes to talk about the “AI arms race” as a battle between labs, the actual battle for the bureaucracy has already been won. The House spending records show that ChatGPT isn’t just a tool for a few tech-savvy interns; it has become the default operating system for a significant chunk of Capitol Hill. This isn’t just about a few subscriptions; it’s about a systemic migration of government labor toward a single proprietary endpoint.
According to TechCrunch, congressional offices are relying on the chatbot for everything from drafting memos and summarizing legislation to handling the endless slog of constituent communications. On the surface, this looks like a win for efficiency. Why spend four hours summarizing a 600-page appropriations bill when a prompt can do it in six seconds? (And let’s be honest, most staffers would rather be doing anything else). But efficiency is a dangerous metric when the output is the law of the land. The real friction isn’t the monthly cost of the Pro plan—which is a rounding error in a federal budget—but the fact that these tools are being used as a substitute for critical reading.
There is a fundamental difference between using an LLM to polish a thank-you note to a donor and using it to synthesize the nuances of a legislative amendment. The former is a clerical task; the latter is a cognitive one. When a staffer asks a model to summarize a bill, they aren’t just saving time—they are outsourcing the actual act of reading and understanding. We are effectively moving toward a system where the people writing the laws are relying on a black box to tell them what they are writing. It turns the legislative process into a game of telephone where the first person in the chain is a probabilistic token-predictor that doesn’t actually know what a “law” is.
It is a bit like a pilot who relies entirely on autopilot without knowing how to actually fly the plane. It works perfectly until the sensors glitch, and suddenly the pilot is staring at a crashing aircraft with no idea which lever to pull. In the context of Congress, the “glitch” isn’t a crash—it’s a hallucination. A subtly misplaced “not” or a fabricated clause in a summary can change the entire meaning of a policy. Or maybe I’m overstating the risk—some might argue that a human always reviews the output—but the records suggest the scale is too large for manual verification to be the primary safety net. In a codebase, you have a compiler to catch errors. In legislation, the “compiler” is the legal system, and the bugs there are called lawsuits and constitutional crises.
Do we really want the people writing the tax code to be treating it like a high school essay prompt? The liability gap here is staggering. If a staffer misses a detail in a bill, it’s human error, and there is a clear chain of accountability. If a model hallucinated a provision and that provision becomes the basis for a committee vote, who is responsible? The OpenAI subscription is paid for by the taxpayer, but the cost of a single high-profile hallucination in a published summary will be far higher than any monthly API fee. We are trading cognitive rigor for a faster turnaround on emails, a trade that looks great on a productivity spreadsheet but terrible in a courtroom.
By Q4 2026, we will see the first formal congressional ethics probe triggered by a hallucinated clause in a published legislative summary. It is inevitable. The sheer volume of usage reported in these spending records means the “statistical certainty” of a major error has already been reached. We are just waiting for the specific document that causes the blow-up. This isn’t a matter of “if” but “when,” as the gap between the model’s confidence and its accuracy is exactly where the most dangerous laws are born.
It is a disaster waiting for a catalyst.