Wenbo Chen
LLM Post-Training & Evaluation for AI Agents
Applied Scientist at Amazon, post-training LLMs for the Amazon Ads agent
I make LLM agents better and prove it: I post-train the models behind Amazon's Ads agent, and I build the benchmarks the field uses to measure agents. My evaluation work includes SkillsBench (NeurIPS 2026), adopted by Meta, Qwen, and Tencent HY3 to evaluate their models, and ClawsBench (COLM 2026). My research also spans reinforcement learning and LLM test-time scaling.
ML PhD from Georgia Tech / NSF AI4OPT with Pascal Van Hentenryck, where my work on scalable data-driven decision making saw major industrial adoption.
- NeurIPS · COLM · ACL · ICML 2026
- SkillsBench 1.8k★
- SkillsBench adopted by Meta, Qwen, Tencent HY3
What I Work On
LLM Post-Training
Post-training LLMs for the Amazon Ads agent (industry, current focus).
- LSFlow (ICML 2026 Spotlight): RL with combinatorial actions
- MARS (first author): 25-47% cheaper test-time scaling
- RL for ride-hailing relocation (JAIR / IJCAI 2023)
- Learning to branch via RL (NeurIPS 2020 Workshop)
LLM & Agent Evaluation
Benchmarks and methods that measure what agents can do and whether they can be trusted.
- SkillsBench (NeurIPS 2026, co-first author): adopted by Meta, Qwen, HY3
- ClawsBench (COLM 2026): productivity-agent capability and safety
- BenchShield: reward-hacking detection for agent benchmarks
- TRIVIA+ (ACL 2026, first author): hallucination detection
Experience
-
Amazon Jun 2024 - PresentApplied Scientist
- LLM post-training for the Amazon Ads agent.
-
Georgia Institute of Technology 2024Ph.D. in Machine Learning, NSF AI Institute AI4OPT
- Learning to optimize, RL, and verification for real-time decisions in power grids and logistics, with industrial adoption.
News
- Sep 2026 SkillsBench accepted at NeurIPS 2026.
- Sep 2026 BenchShield released: we audit 31k+ agent runs on SkillsBench, ClawsBench, and Terminal-Bench 3 for reward hacking, and catch it with 96% accuracy using formal-model-backed instrumentation.
- Aug 2026 Meta evaluates its new open agentic model Muse Glimmer on SkillsBench, joining Qwen and HY3 in adopting it as a core agentic benchmark.
- Jul 2026 ClawsBench accepted at COLM 2026: a benchmark exposing how capable, and how unsafe, LLM productivity agents really are in realistic workspaces.
- Jun 2026 MARS released: a principled risk-controlled early-stopping rule for parallel LLM test-time scaling: 25–47% fewer tokens, full-budget accuracy.
- May 2026 Co-organized the First Workshop on Agent Skills at ACM CAIS 2026, San Jose, CA.
- Apr 2026 Paper on LLM hallucination detection accepted at ACL 2026 (Main Conference).
- Apr 2026 Paper on RL with combinatorial actions accepted at ICML 2026 (Spotlight).
- Feb 2026 SkillsBench released: benchmarking agent skills across diverse tasks.
- Oct 2025 "Boosting Column Generation with GNNs for Joint Rider Trip Planning and Crew Shift Scheduling" published in Transportation Research Part E.
- Jul 2025 "Outbound Load Planning in Parcel Delivery Service Networks" accepted in INFORMS Transportation Science.
- Jun 2024 Joined Amazon as an Applied Scientist, working on LLM Agents and Trustworthy AI.
- May 2024 "Compact Optimality Verification for Optimization Proxies" accepted at ICML 2024.
- Apr 2024 Defended my PhD thesis! News.
- Mar 2024 Research highlighted in AI Magazine (NSF's National AI Institutes).
- Feb 2024 Invited talk at Rice University CMOR Colloquium Series.
- Jan 2024 Won VNN-COMP'23 Award for outstanding benchmark.
- Nov 2023 Awarded Anderson-Interface Fellowship for Excellence in Research (Energy & Sustainable Systems).
- Sep 2023 Two papers accepted in IEEE Transactions on Power Systems. Paper and Paper.
- Jul 2023 One paper accepted in IEEE Transactions on Power Systems. Paper.
- May 2023 Won thesis pitch competition (ML & manufacturing track) at IISE Annual Meeting 2023!
- May 2023 Invited presentation: End-to-End Learning and Optimization thrust for AI4OPT External Advisory Board.
- Apr 2023 Invited presentation: "End-to-End Feasible Optimization Proxies for Large-Scale Economic Dispatch" at INFORMS Annual Meeting 2023.
- Mar 2023 Invited presentation: "Confidence-Aware GNNs for Learning Reliability Assessment Commitments" at IISE Annual Meeting 2023.
- Mar 2023 Featured in NSF AI4OPT newsletter student highlights.
- Jan 2023 Attended 2023 Grid Science Winter School and Conference.
- Oct 2022 Invited presentation: "Learning Optimization Proxies for Large-Scale Security-Constrained Economic Dispatch" at INFORMS Annual Meeting 2022.
Selected Publications
LLM & Agent Evaluation
-
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse TasksNeurIPS 202687 expert tasks with deterministic verifiers and paired with/without-skill runs: curated skills lift agent pass rates by +16.6 pp, and focused skills beat exhaustive ones.
Adopted by Meta Muse Glimmer, Qwen, HY3 -
-
BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation InfrastructurePreprint, 2026Reward-integrity layer for agent benchmarks: static taint analysis finds hacking paths before a run, and runtime evidence detects hacking at 96% accuracy; full-chain recall rises from 23-94% to 77-100% at up to 65% lower cost.
-
Rethinking Evaluation for LLM Hallucination DetectionACL 2026A desiderata for hallucination benchmarks and TRIVIA+, a long-context RAG benchmark with realistic label noise; LLM-as-a-Judge stays competitive.
RL & LLM Reasoning
-
-
Latent Spherical Flow Policy for RL with Combinatorial ActionsICML 2026 (Spotlight)A stochastic flow-matching policy in latent space, with a solver guaranteeing feasible actions: +20.6% over state-of-the-art combinatorial RL.
Spotlight
All publications, including decision making & optimization →