Publishers of data and knowledge need to actively help tech developers ensure AI tools use their information accurately. This includes publishing definitions, limitations, and specific questions alongside their evidence to prevent AI from generating unsupported conclusions or omitting crucial context in search results. Organizations such as national statistics offices publish data that can inform the searches people do online using traditional and, increasingly, AI tools. But this data can be cited in an AI search answer even when the evidence does not support the conclusion.
For AI systems to accurately process information, organizations must provide comprehensive context and ensure consistency across various formats. This means definitions, limitations, and methodologies associated with any dataset or research should be made accessible through the same channels as the data itself. For instance, an employment figure provided by an AI tool should always specify the population, period, provisional status, and a link to its calculation method. Similarly, research summaries must not omit the specific conditions or populations to which findings apply. Furthermore, while organizations are authoritative on statistics, interpreting their causes often requires additional evidence. AI tools should be designed to verify claims against supporting evidence, identify supplementary sources when used, and acknowledge any knowledge gaps. By using a single source to publish both human-readable guidance and machine-readable data, organizations can ensure that updates and corrections are universally applied, maintaining data integrity and historical traceability.
To improve the responsible use of their data by AI, organizations should develop and publish example test questions derived from commonly received inquiries regarding their datasets or studies. These tests should go beyond merely checking for numerical accuracy; they must also verify that AI-generated answers correctly incorporate important limitations (e.g., a study's applicability only to a specific age group) and acknowledge when the available evidence cannot fully answer a question. Such structured evaluation provides developers with practical benchmarks to ensure their AI systems faithfully represent the source material and makes it easier for both publishers and users to identify and address any misinterpretations or failures in AI responses. Developers would also need to extend these tests to unfamiliar questions to assess the system's robustness.
Effective collaboration and a robust feedback mechanism between statistical producers, data redistributors, and AI developers are critical for fostering trust in AI-generated search results. The practice of publishing test questions can serve as a practical foundation for this cooperation, enabling publishers to ensure their sources and limitations are accurately conveyed and understood. Concurrently, developers can leverage these tests to evaluate and refine how their AI systems retrieve and interpret information. A collaborative diagnostic process, involving insights from both publishers (e.g., verifying data dates) and developers (e.g., checking system usage), is essential for pinpointing errors. Additionally, industry-wide cooperation and centralized reporting routes for recurring misinterpretations, particularly organized by sector bodies, would benefit smaller organizations by removing the burden of individual partnerships and ensuring consistent accuracy across AI tools.