Wenbo Chen

Wenbo Chen

LLM Post-Training & Evaluation for AI Agents

Applied Scientist at Amazon, post-training LLMs for the Amazon Ads agent

I make LLM agents better and prove it: I post-train the models behind Amazon's Ads agent, and I build the benchmarks the field uses to measure agents. My evaluation work includes SkillsBench (NeurIPS 2026), adopted by Meta, Qwen, and Tencent HY3 to evaluate their models, and ClawsBench (COLM 2026). My research also spans reinforcement learning and LLM test-time scaling.

ML PhD from Georgia Tech / NSF AI4OPT with Pascal Van Hentenryck, where my work on scalable data-driven decision making saw major industrial adoption.

What I Work On

LLM Post-Training

Post-training LLMs for the Amazon Ads agent (industry, current focus).

LLM Post-Training Reinforcement Learning Test-Time Scaling AI Agents

LLM & Agent Evaluation

Benchmarks and methods that measure what agents can do and whether they can be trusted.

  • SkillsBench (NeurIPS 2026, co-first author): adopted by Meta, Qwen, HY3
  • ClawsBench (COLM 2026): productivity-agent capability and safety
  • BenchShield: reward-hacking detection for agent benchmarks
  • TRIVIA+ (ACL 2026, first author): hallucination detection
Agent Evaluation Benchmarks Agent Skills Reward Hacking

Experience

News

Selected Publications

LLM & Agent Evaluation

RL & LLM Reasoning

All publications, including decision making & optimization →