Autonomous AI research agents can execute long-running scientific workflows, revise code, inspect results, and reuse prior observations. Their reliability depends on memory architectures that retrieve relevant evidence without overwhelming the agent or discarding useful context. This hands-on tutorial presents a low-resource framework for diagnosing, benchmarking, and improving memory systems in autonomous research agents. Participants learn taxonomies of agent memory, trace-based diagnostics for retrieval and utilisation failures, LLM-as-judge evaluation protocols, and standardised benchmarking across heterogeneous memory providers. Exercises use real autonomous-research traces and lightweight open-source tooling that runs on a standard laptop with modest API credits.

Akbar, N.A., Dembani, R., Airlangga, G., Wibowo, R.M., Lenzitti, B., Tegolo, D. (2026). Systematic Diagnosis and Benchmarking of Memory Systems in Autonomous AI Research Agents: A Low-Resource Framework. In KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (pp. 13281-13282) [10.1145/3770855.3816477].

Systematic Diagnosis and Benchmarking of Memory Systems in Autonomous AI Research Agents: A Low-Resource Framework

Akbar, Nur Arifin
;
Lenzitti, Biagio;Tegolo, Domenico
2026-08-08

Abstract

Autonomous AI research agents can execute long-running scientific workflows, revise code, inspect results, and reuse prior observations. Their reliability depends on memory architectures that retrieve relevant evidence without overwhelming the agent or discarding useful context. This hands-on tutorial presents a low-resource framework for diagnosing, benchmarking, and improving memory systems in autonomous research agents. Participants learn taxonomies of agent memory, trace-based diagnostics for retrieval and utilisation failures, LLM-as-judge evaluation protocols, and standardised benchmarking across heterogeneous memory providers. Exercises use real autonomous-research traces and lightweight open-source tooling that runs on a standard laptop with modest API credits.
8-ago-2026
9798400722592
Akbar, N.A., Dembani, R., Airlangga, G., Wibowo, R.M., Lenzitti, B., Tegolo, D. (2026). Systematic Diagnosis and Benchmarking of Memory Systems in Autonomous AI Research Agents: A Low-Resource Framework. In KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (pp. 13281-13282) [10.1145/3770855.3816477].
File in questo prodotto:
File Dimensione Formato  
3770855.3816477.pdf

accesso aperto

Tipologia: Versione Editoriale
Dimensione 826.72 kB
Formato Adobe PDF
826.72 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/713786
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact