Updated 3 hours ago
Agentic flooding: the AI paperwork problem hiding behind government queues

AI Governance

Agentic flooding: the AI paperwork problem hiding behind government queues

A closer reading of 84 cases, complaint statistics and the UK’s Consult rollout shows why AI paperwork needs more than faster drafting.

A queue can grow without more people joining it

*Government Offices Great George Street, London, photographed in 2011. Source: [Wikimedia Commons](https://commons.wikimedia.org/wiki/File:Government_Offices_Great_George_Street.jpg); Carlos Delgado, [CC BY‑SA 3.0](https://creativecommons.org/licenses/by‑sa/3.0/). Landscape display crop.* AI can make a complaint easier to write while making it harder to resolve. A study by Chris Schmitz, Lewis Hammond and Alan Chan documents 84 potential cases of “agentic flooding” across 11 jurisdictions, largely involving people submitting AI‑generated text themselves. Its clearest warning concerns complexity: 76 cases involved more demanding submissions, 50 involved greater volume, and 42 involved both. The August paper cannot establish how much of the change AI caused. The findings, [covered by TechCrunch on September 10](https://techcrunch.com/2026/09/10/ai‑agents‑are‑flooding‑public‑services‑with‑new‑requests/), suggest a sharper question for public services adopting AI: how much work does each additional submission create? [The researchers’ findings](https://arxiv.org/html/2608.16603v2#S4) separate those pressures. Subtracting the study’s overlapping categories produces 34 cases involving complexity alone and eight involving volume alone. Half combined both pressures. These are calculations from the paper’s published totals, not estimates of how common AI overload is across government. The distinction matters because a submission counter cannot show whether a caseworker now needs twice as long to understand each case. An agency could receive the same number of applications and still find its staff falling progressively further behind. Consider a hypothetical team receiving 100 submissions a day. If each requires 20 minutes of work, the incoming workload is about 33 hours. If more elaborate submissions raise the average to 30 minutes, the same inbox now creates 50 hours of work daily. If volume also rises to 120 submissions, the workload reaches 60 hours, an 80% increase from the starting point. These are illustrative numbers, but the arithmetic exposes a practical weakness in counting applications alone: workload depends on both the number of cases and the effort each case requires. That changes what an effective writing assistant should optimise. A useful complaint makes the relevant dates, evidence, disputed decision and requested remedy easy to locate. Adding a longer argument has value only if it helps establish something the decision‑maker needs to know. For teams building such assistants, a worthwhile test would give a reviewer the underlying documents alongside the generated submission, then measure how quickly the reviewer can verify its claims. The output can sound persuasive and still perform poorly on that test. Evidence that is easy to check is a more useful product than paperwork that merely looks formidable. The paper's [original twelve‑service charts](https://arxiv.org/html/2608.16603v2#S1.F1) also show why measures need careful interpretation: the dashed markers identify ChatGPT's release, without establishing that AI caused the changes around them.

A complaint count is not always a count of complaints

The Housing Ombudsman illustrates why the underlying measure deserves attention. Its [Annual Complaints Review](https://www.housing‑ombudsman.org.uk/annual‑complaint‑review‑reports/) records 7,082 determinations between April 2024 and March 2025, up 30% from the preceding year. A determination is a completed decision, so that figure should not be presented as the number of new complaints entering the system. The same official collection describes an earlier increase in complaints alongside poor property conditions, legislative changes, media attention and the inquest into Awaab Ishak’s death. AI is not needed to explain every change in demand. A rising total of completed decisions could reflect greater capacity, clearance of an old backlog, higher incoming demand, or some combination. Establishing which explanation applies requires arrivals, completions and the stock of unresolved cases to be examined together. For readers comparing AI claims across public agencies, that is a useful discipline: first identify what the line on the chart actually counts. “Applications received,” “cases decided” and “people waiting” describe different stages of a service. Treating them as interchangeable can turn a story about processing more work into a story about receiving more work, before the role of AI has even been assessed.

Government is already using AI to read the replies

The response is no longer purely hypothetical. In its methodology for the [Growing up in the online world consultation](https://www.gov.uk/government/consultations/growing‑up‑in‑the‑online‑world‑a‑national‑consultation/outcome/summary‑of‑evidence‑methodology‑and‑organisations‑who‑responded‑to‑the‑consultation‑july‑2026), the UK government describes using Consult to identify themes and count responses mentioning them. Officials reviewed the proposed themes, corrected inaccurate ones and removed irrelevant or repetitive categories. They also reviewed 279 email submissions individually. A further 33,141 campaign‑organised pro forma emails were recorded separately. That is evidence of different treatment for different kinds of material; it does not establish that the campaign emails were AI‑generated. This is a more concrete model for AI‑assisted processing than simply asking a chatbot to summarise an inbox. Theme discovery, counting and closer reading serve different purposes. Repeated messages can communicate the strength of an organised campaign, while a single detailed submission can introduce evidence that changes a policy team’s understanding. A system that compresses everything into the most common themes risks losing that distinction. The operational question is therefore whether staff can still trace a summary back to the responses behind it and identify important material that does not fit a frequent category. There is also a capacity constraint on the remedy itself. The government’s [Consult access page](https://www.gov.uk/government/publications/consult‑register‑your‑interest/consult‑ai‑tool), published in July, says demand exceeds capacity and that departments are joining a waitlist ahead of a full cross‑government rollout planned for 2027. The tool is designed to identify themes and map consultation responses to them, freeing analysts for interpretation. Its availability should therefore be treated as an implementation question, not an assumption that every department can immediately install the same response to rising paperwork. A service facing pressure today still has to decide which work to prioritise with the staff and systems it has.

Reading with AI does not settle who makes the decision

A separate [Ofqual consultation privacy notice](https://www.gov.uk/government/consultations/regulating‑post‑16‑vocational‑and‑technical‑qualifications‑at‑levels‑2‑and‑3/annex‑e‑privacy‑notice‑for‑consultations) makes the division of responsibility explicit. It says tools such as Microsoft Copilot may help identify themes and summarise feedback, with personal and special‑category data anonymised before analysis. Final outcomes remain determined by people. This is a commitment for that consultation, not a description of every government AI deployment, but it demonstrates that the role of the software can be specified much more precisely than “AI handles the responses.” Analysis assistance and authority to decide are separate choices. That separation should shape how public‑sector AI products are judged. A demonstration in which software produces a plausible answer from a sample letter leaves several important questions unanswered: whether it omitted a relevant attachment, whether its summary preserves an unusual objection, and whether a reviewer can recover the original wording. A stronger demonstration would show the complete path from submission to evidence to staff decision, including a case the model cannot confidently interpret. The ability to refer a difficult submission for closer review is part of useful automation; it should be evaluated alongside the speed with which routine material is processed.

The useful metric is a resolved case, not generated text

The UK’s existing [Service Standard](https://www.gov.uk/service‑manual/service‑standard/point‑10‑define‑success‑publish‑performance‑data) offers a practical starting point: define whether a service is solving its intended problem, combine performance data with user research, and collect information across online and offline channels. Applied to AI paperwork, that means looking beyond the number of drafts produced or summaries generated. A department needs to know whether people reach a correct outcome sooner and whether staff spend less time repairing incomplete or misleading information. Faster generation is an intermediate step, not evidence that the whole service has improved. For a measured pilot, incoming submissions, processing time and completed cases should be tracked together, alongside the frequency of requests for clarification and corrections to AI‑assisted work. Those are proposed evaluation measures, not results established by the studies above. They would make it possible to test whether an assistant removes effort or merely transfers it from the applicant to the reviewer. The strongest version of public‑service AI would help someone present a legitimate case clearly and help the agency assess it accurately. Both sides would finish with less work, and the person waiting for a decision would see the benefit.

Share this article

PostShare

Related News