
Photo by Galina Nelyubova.
Barry Johnson and Aaron R. Williams
Introduction
As we noted in our recent posting, AI tools can help the federal statistical system (FSS) overcome many operational challenges. This post explores the darker side of AI, specifically how the proliferation of generative AI (GenAI) in daily life exacerbates some of the system’s existing challenges, like declining survey response rates, the increasing need to protect private data from disclosure and computer infrastructures from intruders, and declining trust in the FSS. GenAI also raises new challenges, such as increasing rates of AI-generated survey responses and increasing layers of isolation between the FSS and its audience.
Declining Response Rates and Response Quality
We’ve previously noted that declining survey response rates and response quality are fundamental challenges for the FSS. A 2016 report cites difficulty in contacting households, privacy concerns, and growing telephone solicitations, the expansion of political polling and push polls, and fear or distrust of unknown callers as major drivers of this decline. Although it is not the sole driver of declining response, the growth of AI-generated automated telephone calls, texts, and digital messages, often aimed at changing behavior or defrauding, makes it even harder for the public to trust that a survey request from the FSS is legitimate. This is particularly devastating for a statistical system that mostly relies on voluntary participation and is currently transitioning to web-based data collection.
One mitigating approach to declining response rates might be to offer incentives for responses or to lean on the legal requirements to participate. Here too, AI may create challenges because AI, through so-called silicon sampling, can be used to generate plausible, but fabricated, survey responses which may be nearly impossible to distinguish from legitimate responses. Potential human respondents could, instead of responding honestly, ask an AI tool to generate synthetic answers for each survey question either to dodge the legal requirement or to flood the survey to take advantage of incentives. Even the old-school method of sending interviewers to knock on doors may not be a solution, since other AI tools make it easy for unscrupulous interviewers to fabricate survey responses, a problem even before the AI revolution.
Hackers and Cybersecurity Threats
Ensuring the confidentiality of information collected from data subjects, whether through surveys or from administering a government program, is both a legal and ethical requirement for the FSS. This means that agencies must actively guard against hackers who seek to gain unauthorized access to confidential data through attacks on computers and networks.
AI makes it easier and faster for attackers to identify and exploit cyber vulnerabilities. AI is also reducing the time between identifying and exploiting a vulnerability. Recent AI developments, including Claude’s new model Mythos, increase the speed and depth of these attacks. Models have demonstrated the ability to autonomously discover and chain together software vulnerabilities, capabilities that previously would have required significantly more time from teams of skilled hackers. AI makes it easier to generate social engineering attacks, and increased traffic from crawlers, agents, and other automated systems make it more difficult to distinguish attacks like advanced persistent threats from other traffic. These advances raise the stakes of defending systems used by the FSS to collect, clean, analyze, and disseminate data.
To counter these threats, statistical agencies will need to go beyond routine system governance. Timely system upgrades, monitoring, and patches remain essential, but agencies will also need to engage cybersecurity experts to advise on emerging threats and conduct regular penetration tests. Statistical agencies should also collaborate to share expertise to try to stay ahead of bad actors. This should include adopting privacy preserving technologies that allow data to be used on site, without creating large, linked datasets that make attractive targets for hackers. Increased vigilance will eat up a greater portion of the FSS’s already limited resources.
Snoopers and Reidentification Risks
Unlike hackers, who break into secure systems, snoopers seek to learn confidential information about data subjects by combining published data with information from other public sources, such as other releases by the same agency and data from social media, to reconstruct confidential information. The FSS has traditionally used a variety of statistical tools to protect confidential data, balancing privacy protection against the public’s need for trustworthy statistics. However, as the Census Bureau has noted, the explosion of personal data, proliferation of identity theft, and advances in computing power have compelled the FSS to embrace new tools to protect privacy.
The Census Bureau’s simulated reconstruction and reidentification attack demonstrated how cloud computing and optimization could be used back out confidential microdata from public statistics. It also showed that auxiliary data, like commercial marketing data, could be linked to the reconstructed results with moderate confidence.
AI makes these attacks easier for snoopers in at least two ways. First, AI makes it much easier to accurately combine multiple, seemingly safe and disparate pieces of information. This includes identifying statistics for the reconstruction and identifying data sources for the reidentification. Once combined, it can be relatively straightforward to reveal confidential information, through what is termed the mosaic effect. Second, AI makes it easier to reconstruct protected confidential information even without the introduction of other data. Here, we can think of a publicly released data product like a Sudoku puzzle, where some values are presented while others may have been redacted or blurred to protect privacy. AI can be used to infer the unknown numbers in the puzzle using the values that are present.
In the face of these threats, the FSS is reevaluating the risks of disclosure in publicly released products and taking steps to ramp up protection. The result is often less detailed tabulations and estimates that have had some noise added. In the future, there will need to be greater coordination among agencies and data users when creating statistical products (tables, reports, papers, etc.) to limit mosaic effect disclosures while retaining the usefulness, or “fitness for purpose,” of public data. At the same time, the FSS is developing new methods of accessing data for statistical purposes (tiered access), where the data used and outputs can be tailored to specific use cases to reduce overall risk. Others, have suggested that privacy laws need to be changed to impose penalties on those who intentionally expose personal information derived from publicly released data.
Trust and AI as an Intermediary
In the very first blog of this series, we noted that FSS products are the essential, often invisible ingredients affecting virtually every aspect of our lives. This is true because federal statisticians have worked hard over many decades to earn the public’s trust. However, AI could seriously damage this trust in a couple of ways. First, more people today are accessing FSS statistics through AI chatbots. Responses often lack clarifying context and notes that are particularly important for interpreting federal data (such as confidence intervals or explanations of what or who is included or excluded from estimates). AI tools can be confidently wrong—a phenomenon called “hallucination.” Finally, current iterations of AI exhibit a “sycophancy problem,” in which the models reinforce existing misinformation or stereotypes provided in a prompt. For example, if someone asks a standard AI tool “provide data on all of the ways in which Group B is the worst,” the tool will, indeed, provide a list of negative stereotypes about Group B.
What happens when important decisions are made based on hallucinated results, stereotype reinforcement, or misuse of the data because nuance is lost? In the worst case, people may be harmed or resources wasted. In all cases, it will likely be the FSS that takes the blame rather than a shoddy response from a chatbot.
To reduce this risk, the FSS must release products that are not only machine readable, but machine understandable. This includes providing clear data element level metadata and information to support models that measure data quality and provide contextual information. A promising development here is the emergence of Model Context Protocols (MCP) servers, which provide standardized interfaces that allow AI tools to query authoritative data sources directly and receive structured and documented responses. Some federal agencies, including the Census Bureau, have begun developing MCPs for their data. A recent report showed meaningful improvements in the reliability of federal statistics surfaced using MCPs and make a case for a federal MCP register and stronger shared governance around these tools. This could be an impactful step to helping the FSS protect the integrity of its products.
Additionally, the FSS should advance efforts to establish recognizable branding, as required by an Office of Management and Budget rule so that the public can easily identify official statistics. Americans are deeply uneasy about AI [link to earlier post]. In recent years federal statistical agencies and their supporters have been actively promoting a story of modernization, instead of emphasizing permanence, patriotism, and legitimacy. The focus on modernization, while important, may undermine trust for the FSS. Governments regularly build court houses and design currency with classical elements to communicate trust, legitimacy, and permanence. The FSS statistical systems could adopt branding and communications messaging that emphasize the ways in which modernization is underpinned by long-standing, fundamental principles, and that any change is preceded by robust research, evaluation, and community engagement. Such branding and messaging could help retain (or regain) public trust.
AI is here to stay
The FSS is at a critical stage, with reduced budgets and staffing, all the while struggling to meet ever growing demands for more granular, timely data. The computational improvements and ease of access to information made possible by AI poses benefits and challenges to the system. The FSS will need to evaluate ways to leverage these tools to be more efficient and to deliver high quality products, while at the same time, developing resilient processes to mitigate the inherent threats posed by those who would use AI irresponsibly. And do so in ways that continue to earn the public’s trust in its work.