00 /

News

Availability. I am applying for PhD positions in the current (2026–2027) admissions cycle, and am also seeking internship opportunities for the term after next, as I will graduate after the coming semester. Feel free to reach out or connect; I'd be glad to chat.

01 /

About

I'm a junior at Johns Hopkins University majoring in Applied Mathematics and Statistics with a minor in Computer Science. I work with Prof. Daniel Khashabi in the Intelligence Amplification Lab at the Center for Language and Speech Processing, and am co-supervised by Prof. Tianmin Shu, director of the Social Cognitive AI (SCAI) Lab. I'm currently also interning with Prof. Andrea Zanette.

My research is in natural language processing and machine learning, with a focus on understanding and improving the reasoning behavior of large language models — including self-correction and self-improvement, harness evolution, confidence calibration, efficient reasoning, memory, and creativity. I believe an intelligent reasoning agent or system is an aggregation of many diverse abilities, and building such a robust system requires familiarity with several interconnected research directions.

02 /

Research Interests

Broadly: NLP, machine learning, and deep learning. More specifically:

  • 01

    Improving LLM reasoning — particularly self-correction and incorporation of external feedback.

  • 02

    Confidence calibration, hidden-state probing, and steering for reliable reasoning.

  • 03

    Efficient reasoning and inference-time scaling — when, and how much, to think.

  • 04

    Evaluating and improving LLM creativity, and using LLMs as judges.

  • 05

    Mathematical reasoning, robustness under question perturbation, and weak-to-strong generalization.

  • 06

    LLM agents in open-ended environments.

03 /

Publications

* indicates equal contribution.

NeurIPS 2025

Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback

Dongwei Jiang*, Alvin Zhang*, ··· , Daniel Khashabi

ICML 2026 · NeurIPS 2025 Workshop on Efficient Reasoning

Compute When Worth It: Risk Control for Reasoning on a Compute Budget

Anushri Suresh*, Alvin Zhang*, ··· , Daniel Khashabi

ICML 2026

Trust Functions: Near Lossless Weak-to-Strong Generalization by Learning to Trust the Weak Teacher

Alvin Zhang*, Arda Uzunoglu*, Daniel Khashabi

TMLR

CreativityPrism: A Holistic Benchmark for Machine Creativity

Zhaoyi Joey Hou, Alvin Zhang, ··· , Daniel Khashabi, Xiang Lorraine Li

Under review

AgentOdyssey: Open-Ended Text Game Generation for Test-Time Continual Learning Agents

Zheyuan Zhang, Zehao Wen, Alvin Zhang, Andrew Wang, Jianwen Xie, Daniel Khashabi, Tianmin Shu

04 /

Blog

  • Jul 2026 Jacobian counterexample playground. An interactive note on Alpöge’s counterexample to the Jacobian conjecture: compose shears and rescalings over the det-1 map, track an exact determinant ledger, and watch the three-point collision no composition can separate.
  • Jul 2026 Loop deeper, or adapt? Test-time training in looped transformers. Test-time training in looped / recurrent-depth transformers (Ouro): where and when a tiny RMSNorm-scale update helps, and why most of the four-shot GSM8K gain is reusable prompt calibration.
05 /

Education

Johns Hopkins University  ·  Baltimore, MD

B.S. in Applied Mathematics and Statistics  ·  Minor in Computer Science
GPA: 4.0 / 4.0  ·  Dean's List every semester

Selected coursework: Honors Mathematical Statistics, Bayesian Statistics, Optimization, Time Series Analysis, Machine Learning: Deep Learning, Natural Language Processing, Self-Supervised Models, AI Agents, Honors Discrete Mathematics, Computational Cognitive Neuroscience of Vision, AI and Democracy.

06 /

Experience

Research Intern

Advisor: Prof. Andrea Zanette

Undergraduate Researcher

Johns Hopkins University  ·  Advisors: Prof. Daniel Khashabi (Intelligence Amplification Lab, CLSP) and Prof. Tianmin Shu (Social Cognitive AI Lab)

Working across several projects on LLM self-correction and feedback incorporation, risk-controlled early stopping for reasoning models, weak-to-strong generalization, creativity evaluation, and agent learning in open-ended environments.

Software Development Engineer Intern

Caterpillar Inc.  ·  Wuxi, Jiangsu, China

Developed recurrent neural networks to simulate current and voltage in electrical motors, exploring neural-network replacements for traditional simulation pipelines.

07 /

Teaching

Course Assistant  ·  EN 601.465 Natural Language Processing

Instructor: Prof. Jason Eisner

Undergraduate TA  ·  EN 553.361 Introduction to Optimization

Instructor: Prof. Donniell Fishkind

Led TA sessions, ran problem-solving recitations, and graded homework and exams.

08 /

Skills

Languages
English  ·  Chinese (Mandarin)
Programming
Python  ·  Java  ·  MATLAB  ·  R  ·  Mathematica
ML / NLP
PyTorch  ·  Transformers  ·  VeRL  ·  vLLM  ·  SGLang
Systems
Linux  ·  HPC clusters
Areas
Machine Learning  ·  Deep Learning  ·  Reinforcement Learning