Andy Zhou

Andy Zhou

I'm a co-founder of Intology, a research lab developing Artificial Scientists: systems that read the literature, form hypotheses, run experiments, and communicate what they discover.

Our work asks a basic question: what would it take for an AI system to make a genuine scientific contribution? Zochi was the first AI system to pass peer review at a top-tier venue, and Locus is the first to outperform human experts at AI R&D.

Before Intology I did research at Meta and Virtue AI. I am currently on-leave from my PhD program at Carnegie Mellon University where I was awarded the NSF Graduate Research Fellowship.

01

Research program

Artificial scientists, autonomous experimentation, and evaluations of genuine research ability.

  1. Artificial Scientist for AI R&D

    Locus

    2025 —

    A long-horizon research system that plans and steers many experiments in parallel over multi-day horizons. Locus was the first AI system to beat human experts on RE-Bench at equal time and compute; it now leads PostTrainBench, surpasses the official human-tuned Qwen3-1.7B instruct checkpoint under PostTrainBench+, and has post-trained a model deployed in production.

    Selected results

    • 44.7 PostTrainBench SOTA (verified)
    • 51.6 on PostTrainBench+, above Qwen3-1.7B-Instruct
    • 4th among accounts entered in all live prize-money Kaggle competitions
  2. The first Artificial Scientist

    Zochi

    2025

    Zochi takes a research question from literature review through experimentation to a written paper. Its paper Tempest was accepted to the main proceedings of ACL 2025 — the first fully AI-generated discovery to clear peer review at an A*-ranked venue, scoring in the top 8.2% of submissions by meta-review score.

    Selected results

    • ACL 2025 main proceedings
    • Top 8.2% of submissions
    • Peer-reviewed papers at multiple ICLR 2025 workshops
  3. Measuring real research ability

    NanoGPT-Bench

    2026

    Benchmarks for research agents are easy to saturate and easy to contaminate. NanoGPT-Bench drops agents into the GPT-2 pretraining speedrun at a fixed human world record and asks them to make it faster, with no internet and no human help. Frontier coding agents recover less than 10% of what the human community achieved over the following five months, spending most of their compute on hyperparameter tuning instead of the algorithmic work that actually moves the record.

    Selected results

    • 9.3% of human progress recovered, at best
    • 512 H100-hours per agent

02

Selected publications

Agents, robustness, and AI safety. The complete list is on Google Scholar.

  1. Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

    ICML 2024

    Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, Yu-Xiong Wang

    The first search algorithm for language agents, combining reasoning, acting, and planning into one tree search. Now implemented in LangChain and LlamaIndex.

  2. Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

    NeurIPS 2024 Spotlight

    Andy Zhou, Bo Li, Haohan Wang

    A defense objective and an algorithm for finding trigger tokens that enforce harmless behavior, transferring across jailbreaks and models.

  3. AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

    arXiv 2024

    Yi Zeng*, Yu Yang*, Andy Zhou*, Jeffrey Tan*, Yuheng Tu*, Yifan Mai*, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li

    5,694 prompts spanning 314 risk categories drawn from 16 company policies and 8 government regulations.

  4. KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data

    ACM CCS 2024

    Andy Zhou, Xiaojun Xu, Ramesh Raghunathan, Alok Lal, Xinze Guan, Bin Yu, Bo Li

    Domain knowledge as logical constraints over an ensemble of specialized models, deployed for anomaly detection on production graph data.

  5. Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models

    NeurIPS 2023

    Andy Zhou, Jindong Wang, Haohan Wang, Yu-Xiong Wang

    Teacher gradients generate hard augmentations during distillation, yielding the most OOD-robust ResNet34 and ResNet50 of their time.

All publications (14)
  1. RedCode: Risky Code Execution and Generation Benchmark for Code Agents

    NeurIPS Datasets & Benchmarks 2024

    Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, Bo Li

  2. Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

    NeurIPS 2024

    Haibo Jin, Andy Zhou, Joe D. Menke, Haohan Wang

  3. Tamper-Resistant Safeguards for Open-Weight LLMs

    arXiv 2024

    Rishub Tamirisa*, Bhrugu Bharathi*, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, Mantas Mazeika

  4. AI Risk Categorization Decoded: From Corporate Policies to Government Regulations

    GenAI & Law @ ICML 2024

    Yi Zeng*, Kevin Klyman*, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li

  5. Towards Robust Unlearning in LLMs

    SeT LLM @ ICLR 2024

    Rishub Tamirisa, Bhrugu Bharathi, Andy Zhou, Bo Li, Mantas Mazeika

  6. FedSelect: Personalized Federated Learning with Customized Selection of Parameters for Fine-Tuning

    CVPR 2024

    Rishub Tamirisa, Chulin Xie, Wenxuan Bao, Andy Zhou, Ron Arel, Aviv Shamsian

  7. GUARD: Role-playing to Generate Natural-Language Jailbreakings to Test Guideline Adherence of Large Language Models

    SeT LLM @ ICLR 2024

    Haibo Jin*, Ruoxi Chen*, Andy Zhou, Jinyin Chen, Yang Zhang, Haohan Wang

  8. YouTubePD: A Multimodal Benchmark for Parkinson's Disease Analysis

    NeurIPS Datasets & Benchmarks 2023

    Andy Zhou*, Samuel Li*, Pranav Sriram*, Xiang Li*, Jiahua Dong*, Ansh Sharma, Yuanyi Zhong, Shirui Luo, Volodymyr Kindratenko, George Heintz, Christopher Zallek, Yu-Xiong Wang

  9. A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance

    ICCV 2023

    Zeyi Huang, Andy Zhou, Zijian Lin, Mu Cai, Haohan Wang, Yong Jae Lee

03

Writing from Intology

Research notes, technical reports, and results from the lab.

  1. Aug 2026

    Scaling Automated Post-Training

    Locus leads PostTrainBench, surpasses human-tuned models under PostTrainBench+, and post-trains a production model for Bubble.

  2. May 2026

    Can AI Agents Actually Do Research?

    NanoGPT-Bench measures how much of five months of human research progress frontier agents can recover under controlled conditions.

  3. Nov 2025

    Previewing Locus

    Introducing the first AI system to outperform human experts at AI R&D.

  4. May 2025

    Zochi Publishes at ACL 2025

    The first AI system to independently make a discovery and pass peer review at an A*-ranked scientific venue.

  5. Mar 2025

    The Zochi Technical Report

    How an autonomous AI system conducts the scientific process end to end.

Read all Intology writing