Blog icon

As artificial intelligence becomes more capable, it is increasingly being used to support decisions in areas such as healthcare, cybersecurity and finance. While most processes for evaluating AI focus on whether a system produces the correct answer, researchers are beginning to recognise that accuracy alone may not provide a complete picture of an AI system's trustworthiness.

For example, an AI system may return the right answer while also shaping how that answer is presented, and in doing so, influence how a person or another AI system processes that answer.

In partnership with the Australian AI Safety Institute, CSIRO is exploring this challenge through research in scalable AI oversight. Because understanding the broader effects of AI is becoming an important part of ensuring systems remain safe, reliable and trustworthy.

The Epistemic Safety in Scalable Oversight report

The Epistemic Safety in Scalable Oversight report introduces a new framework for evaluating AI systems that looks beyond whether they produce the correct answer to also examine how they communicate information and whether those communications can influence the judgement of whoever is overseeing the system, whether that's a person or another AI system.

The research treats AI oversight as both a verification and communication challenge, introducing new concepts and measurement techniques to investigate potential influence during AI interactions.

The report provides a foundation for more robust AI assurance and safety research, helping researchers, policymakers and industry better understand the trustworthiness of increasingly capable AI systems and supporting the development of safer, more accountable AI.

Potential applications

This research has potential relevance across a broad range of areas, including

  • AI safety and alignment research
  • AI assurance and evaluation
  • Government and regulatory oversight
  • Standards development
  • Security and other high-stakes decision-making environments
  • Future human-AI and AI-to-AI oversight systems

Access the report

Explore the full research report, benchmark datasets, evaluation tools and supporting materials to learn more about the framework and its findings.

This project was undertaken in collaboration with the Australian AI Safety Institute (AISI) and supported by funding from the Department of Industry, Science and Resources (DISR), as part of research into AI alignment tools and techniques.

As artificial intelligence becomes more capable, it is increasingly being used to support decisions in areas such as healthcare, cybersecurity and finance. While most processes for evaluating AI focus on whether a system produces the correct answer, researchers are beginning to recognise that accuracy alone may not provide a complete picture of an AI system's trustworthiness.

For example, an AI system may return the right answer while also shaping how that answer is presented, and in doing so, influence how a person or another AI system processes that answer.

In partnership with the Australian AI Safety Institute, CSIRO is exploring this challenge through research in scalable AI oversight. Because understanding the broader effects of AI is becoming an important part of ensuring systems remain safe, reliable and trustworthy.

The Epistemic Safety in Scalable Oversight report

The Epistemic Safety in Scalable Oversight report introduces a new framework for evaluating AI systems that looks beyond whether they produce the correct answer to also examine how they communicate information and whether those communications can influence the judgement of whoever is overseeing the system, whether that's a person or another AI system.

The research treats AI oversight as both a verification and communication challenge, introducing new concepts and measurement techniques to investigate potential influence during AI interactions.

The report provides a foundation for more robust AI assurance and safety research, helping researchers, policymakers and industry better understand the trustworthiness of increasingly capable AI systems and supporting the development of safer, more accountable AI.

Potential applications

This research has potential relevance across a broad range of areas, including

  • AI safety and alignment research
  • AI assurance and evaluation
  • Government and regulatory oversight
  • Standards development
  • Security and other high-stakes decision-making environments
  • Future human-AI and AI-to-AI oversight systems

Access the report

Explore the full research report, benchmark datasets, evaluation tools and supporting materials to learn more about the framework and its findings.

This project was undertaken in collaboration with the Australian AI Safety Institute (AISI) and supported by funding from the Department of Industry, Science and Resources (DISR), as part of research into AI alignment tools and techniques.

Collaborate with us

CSIRO welcomes opportunities to work with researchers, industry partners, evaluators and policymakers interested in advancing AI oversight, assurance and safety.

If you would like to discuss potential collaborations or future research opportunities, we'd love to hear from you.

Contact us now