Andy Zhou

San Francisco · Intology

Andy Zhou

I build AI systems that do science.

I'm a co-founder of Intology, a research lab automating the process of discovery. We build Artificial Scientists: systems that read the literature, form a hypothesis, run the experiments, and write up what they found — with no human in the loop.

Two exist so far. Zochi was the first AI system to pass peer review at a top-tier venue. Locus is the first to outperform human experts at AI R&D.

Before Intology I did research at Meta and Virtue AI, and founded Lapis Labs, a student-led group that published 20+ papers with collaborators including Google DeepMind, the Center for AI Safety, and IBM Research.

Portrait of Andy Zhou

What I'm building

Locus

Artificial Scientist for AI R&D

2025 —

A long-horizon research system that runs its own experiments for days at a time. Locus is the first AI system to beat human experts on RE-Bench given the same time and compute, and in January 2026 it set a world record on the human NanoGPT Speedrun leaderboard with a fused Triton kernel it designed, implemented, and debugged itself.

  • 1.30 vs 1.27 human expert on RE-Bench
  • SOTA on KernelBench and MLE-Bench Lite
  • 64-hour autonomous runs

Zochi

The first Artificial Scientist

2025

Zochi takes a research question from literature review through experimentation to a written paper. Its paper Tempest was accepted to the main proceedings of ACL 2025 — the first fully AI-generated discovery to clear peer review at an A*-ranked venue, scoring in the top 8.2% of submissions by meta-review score.

  • ACL 2025 main proceedings
  • Top 8.2% of submissions
  • Peer-reviewed papers at multiple ICLR 2025 workshops

NanoGPT-Bench

Measuring real research ability

2026

Benchmarks for research agents are easy to saturate and easy to contaminate. NanoGPT-Bench drops agents into the GPT-2 pretraining speedrun at a fixed human world record and asks them to make it faster, with no internet and no human help. Frontier coding agents recover less than 10% of what the human community achieved over the following five months, spending most of their compute on hyperparameter tuning instead of the algorithmic work that actually moves the record.

  • 9.3% of human progress recovered, at best
  • 512 H100-hours per agent

Selected research

Agents, robustness, and AI safety, from before Intology. The complete list lives on Google Scholar.

  1. Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

    ICML 2024

    Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, Yu-Xiong Wang

    The first search algorithm for language agents, combining reasoning, acting, and planning into one tree search. Now implemented in LangChain and LlamaIndex.

  2. Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

    NeurIPS 2024 Spotlight

    Andy Zhou, Bo Li, Haohan Wang

    A defense objective and an algorithm for finding trigger tokens that enforce harmless behavior, transferring across jailbreaks and models.

  3. AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

    arXiv 2024

    Yi Zeng*, Yu Yang*, Andy Zhou*, Jeffrey Tan*, Yuheng Tu*, Yifan Mai*, Kevin Klyman, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li

    5,694 prompts spanning 314 risk categories drawn from 16 company policies and 8 government regulations.

  4. KnowGraph: Knowledge-Enabled Anomaly Detection via Logical Reasoning on Graph Data

    ACM CCS 2024

    Andy Zhou, Xiaojun Xu, Ramesh Raghunathan, Alok Lal, Xinze Guan, Bin Yu, Bo Li

    Domain knowledge as logical constraints over an ensemble of specialized models, deployed for anomaly detection on production graph data.

  5. Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models

    NeurIPS 2023

    Andy Zhou, Jindong Wang, Haohan Wang, Yu-Xiong Wang

    Teacher gradients generate hard augmentations during distillation, yielding the most OOD-robust ResNet34 and ResNet50 of their time.

All publications (14)
  1. RedCode: Risky Code Execution and Generation Benchmark for Code Agents

    NeurIPS Datasets & Benchmarks 2024

    Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, Bo Li

  2. Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

    NeurIPS 2024

    Haibo Jin, Andy Zhou, Joe D. Menke, Haohan Wang

  3. Tamper-Resistant Safeguards for Open-Weight LLMs

    arXiv 2024

    Rishub Tamirisa*, Bhrugu Bharathi*, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, Mantas Mazeika

  4. AI Risk Categorization Decoded: From Corporate Policies to Government Regulations

    GenAI & Law @ ICML 2024

    Yi Zeng*, Kevin Klyman*, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li

  5. Towards Robust Unlearning in LLMs

    SeT LLM @ ICLR 2024

    Rishub Tamirisa, Bhrugu Bharathi, Andy Zhou, Bo Li, Mantas Mazeika

  6. FedSelect: Personalized Federated Learning with Customized Selection of Parameters for Fine-Tuning

    CVPR 2024

    Rishub Tamirisa, Chulin Xie, Wenxuan Bao, Andy Zhou, Ron Arel, Aviv Shamsian

  7. GUARD: Role-playing to Generate Natural-Language Jailbreakings to Test Guideline Adherence of Large Language Models

    SeT LLM @ ICLR 2024

    Haibo Jin*, Ruoxi Chen*, Andy Zhou, Jinyin Chen, Yang Zhang, Haohan Wang

  8. YouTubePD: A Multimodal Benchmark for Parkinson's Disease Analysis

    NeurIPS Datasets & Benchmarks 2023

    Andy Zhou*, Samuel Li*, Pranav Sriram*, Xiang Li*, Jiahua Dong*, Ansh Sharma, Yuanyi Zhong, Shirui Luo, Volodymyr Kindratenko, George Heintz, Christopher Zallek, Yu-Xiong Wang

  9. A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance

    ICCV 2023

    Zeyi Huang, Andy Zhou, Zijian Lin, Mu Cai, Haohan Wang, Yong Jae Lee

Updates

  1. May 2026

    We released NanoGPT-Bench, an evaluation of whether coding agents can actually do research. The best one recovers 9.3% of five months of human progress.

  2. Jan 2026

    Locus set a world record on the human NanoGPT Speedrun leaderboard with a fused Triton kernel it designed on its own. Later human records built on top of it.

  3. Nov 2025

    We introduced Locus, the first AI system to outperform human experts at AI R&D.

  4. May 2025

    Zochi published at ACL 2025 Main — the first AI system to independently pass peer review at an A* venue.

  5. Mar 2025

    We launched Intology and published the Zochi technical report.

  6. Sep 2024

    Three papers at NeurIPS 2024, with RPO accepted as a Spotlight (top 3%), and our AIR-Bench work was covered by Wired.

  7. May 2024

    LATS was accepted at ICML 2024 and is now implemented in LangChain and LlamaIndex.