The ATLAS test harness: Evaluating AI systems for humanities research

Abstract representation of Large Language Models through train tracks
Map

Share via

More Information

MDAP

mdap-info@unimelb.edu.au

  • Seminar

With Prof James Smithies, Director, HASS Digital Research Hub, ANU College of Arts and Social Sciences

The prevailing discourse on Large Language Models (LLMs) suggests that humanities research, like many other fields, is about to be shaped by intellectual automation. This presents complex ethical and technical challenges that need to be embraced.

A similar convergence of technological potential and hype prompted the expansion of digital humanities in the early twenty-first century. The difference now is:

  • We have a much better understanding of the epistemological, methodological, and ideological implications of digital technologies
  • There is a sense that AI technology is too complex, expensive, and opaque for humanities researchers to be anything other than passive consumers.

This creates a contradictory socio-technical context it is essential to break out of – using techniques of critical technical practice.

This talk presents ATLAS, a key output of the AI as Infrastructure (AIINFRA) project, a transnational collaboration spanning Australia, Aotearoa New Zealand, and the United Kingdom. ATLAS serves as a ‘test harness’ for conducting reproducible and transparent experiments with multiple LLMs and text corpora, built with open-source code.

The prototype is designed to help researchers ‘look under the hood’ of LLM-based tools and services, revealing the many design choices and configurations that shape their output. This includes detailed technical metadata about each experiment’s ‘calibration’ – such as model version, word embeddings, vector store characteristics, and system prompts – alongside source documents informing the foundation model’s responses.

The goal of ATLAS was not to produce another AI product for non-technical users, but to expose LLM Retrieval Augmented Generation (RAG) to transparent and holistic analysis, and demonstrate the continued value of open source tinkering and experimentation.