About me

Hi👋! I am a fourth-year PhD student in Computer and Information Science at the University of Pennsylvania, advised by Prof. Dan Roth, and currently a research intern at Meta. I received my Master’s degree in Economics and Computer Science from Duke University, advised by Prof. Sam Wiseman. Before that, I graduated from Renmin University of China (RUC) with a major in Mathematics and Applied Mathematics and a minor in Computer Science, where I worked with Prof. Jing Zhang and Prof. Xin Zhao.

My research goal is to build LLM systems that are reliable:

  • Reliable reasoning with symbolic tools. I combine the flexibility of LLMs with the reliability and verifiability of symbolic tools (solvers, probabilistic programs, and formal verification) so that systems reason, plan, and make decisions consistently. Examples include Bayesian inference for trustworthy decision probabilities (BIRD), solver-based verification of chain-of-thought (VeriCoT), treating chain-of-thought as tractable probabilistic programs (Copper), multi-agent uncertainty estimation for black-box LLMs (DiverseAgentEntropy), and showing that the gains from tool-augmented reasoning come mainly from reliable execution rather than from writing reasoning as code (Is Code Better Than Language?).

  • Reliable agents. I study agents that proactively gather information, reason over evidence, and act safely under uncertainty. This includes training agents to discover reusable abstractions, skills, and computational structure instead of memorizing task-specific solutions(ReuseRL), diagnosing search agents’ process through evidential query graphs (SearchAtlas), evaluating whether agents recognize risks and act on them (AURA-Eval).

đź“‘ Selected Research Projects

CoTs as Tractable Probabilistic Programs
Kyle Richardson *, Yu Feng *, Poorva Garg, Junyan Cheng, Guy Van den Broeck, Dan Roth
NeurIPS 2026; * equal contribution; earlier version at the 9th Workshop on Tractable Probabilistic Modeling (TPM 2026)

VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
Yu Feng, Nathaniel Weir, Kaj Bostrom, Sam Bayless, Darion Cassel, Sapana Chaudhary, Benjamin Kiesl-Reiter, Huzefa Rangwala
ICLR 2026

BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models
Yu Feng, Ben Zhou, Weidong Lin, Dan Roth
ICLR 2025 (Oral)

Skill Reuse as Compression in Agentic RL
Zhikun Xu, Yu Feng, Jacob Dineen, Taiwei Shi, Jieyu Zhao, Ben Zhou
EMNLP 2026 (Main)

Rethinking LLM Uncertainty: A Multi-Agent Approach to Estimating Black-Box Model Uncertainty
Yu Feng, Phu Mon Htut, Zheng Qi, Wei Xiao, Manuel Mager, Nikolaos Pappas, Kishaloy Halder, Yang Li, Yassine Benajiba, Dan Roth
EMNLP 2025 (Findings)

BLINK: Multimodal Large Language Models Can See but Not Perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna
ECCV 2024