- Formulated a unified Gap Detection framework connecting three forms of AI content homogenization — LLM output diversity, GEO market positioning, and synthetic training-data collapse — via Demand × Defensibility optimization in embedding space.
- Developed the Unmet-Demand Theorem and a Nash-product objective; reduced LLM response collision by 12.8% on curated topics and 10.8% across hundreds of INFINITY-CHAT responses.
Xucheng Yu
Trustworthy AI • LLM Safety and Reliability • Agentic and Multi-Agent Systems
M.Eng. in Electrical & Computer Engineering, University of Illinois Urbana-Champaign
xy63@illinois.edu · +86 187 6288 3800 · Urbana-Champaign, IL
About
I am a researcher working on trustworthy foundation models. My work centers on adversarial robustness, auditable reasoning and evaluation, and reliable agentic and multi-agent decision-making under uncertainty. I am drawn to the questions of how large language models behave when they are pushed, probed, and deployed at scale — and how we can make them safer and more dependable in the process.
I recently completed my Master of Engineering in Electrical and Computer Engineering at the University of Illinois Urbana-Champaign (GPA 4.0/4.0), where I work with Dr. Haohan Wang and Dr. Huan Zhang. Before UIUC, I earned dual B.Sc. degrees in Mathematics with Computer Science from the Guangdong Technion — Israel Institute of Technology (GTIIT) and the Technion — Israel Institute of Technology, graduating on the Dean's List. Alongside research, I build production systems — from GEO auditing platforms to natural-language tool orchestration.
News
- Sep 2026New preprint MonitorBench accepted at COLM 2026.
- Sep 2026Submitted Gap Detection for GEO Positioning and LLM Output Diversity to ICLR 2027 and Rhetorical Distortion Detection across Platforms to WWW 2027.
- Aug 2026Understanding Content Moderation through Restricted Books accepted at AIES 2026.
- Aug 2026Prompt Stability Matters accepted at CPAL 2026.
- May 2026SCI-Defense and HEART submitted to NeurIPS 2026; preprint SCI-Defense available on arXiv.
- May 2026Joined TopCited.ai and Heelo as a Research Engineer.
Publications & Manuscripts
Author names in bold indicate my contribution. Submitted manuscripts are under peer review.
-
SCI-Defense: Defending Manipulation Attacks from Generative Engine OptimizationSubmitted to NeurIPS 2026 arXiv:2605.21948
-
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language ModelsAccepted at COLM 2026
-
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to WarningAccepted at AIES 2026
-
HEART: Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool PrimitivesSubmitted to NeurIPS 2026
-
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose SystemsAccepted at CPAL 2026
-
Closed-Loop Self-Improving Scientific Paper Generation: An Iterative Framework Combining AI Generation, Automated Review, and Targeted OptimizationManuscript
-
Gap Detection for GEO Positioning and LLM Output DiversitySubmitted to ICLR 2027
-
Rhetorical Distortion Detection across PlatformsSubmitted to The Web Conference (WWW) 2027
-
GEO Survey and Unified Evaluation BenchmarkSubmitted to IEEE TKDE
Research Experience
Research Assistant, University of Illinois Urbana-Champaign
- Designed a multi-platform, multi-language classifier for account-level rhetorical distortion across 332 accounts and 13,219 posts from RSS/Newsletter, YouTube, Bluesky, Reddit, Weibo, and Twitter/X.
- Built a three-tier pipeline (rule-based signals, negative-pattern filtering, GPT-4o-mini verification) over five distortion dimensions; validated on 200 human-annotated posts with Cohen's κ = 0.76 and a 12-point precision gain over a rule-only baseline.
- Systematically characterized adversarial manipulation of LLM ranking systems across 20 GEO pipelines, establishing a benchmark foundation for evaluating trustworthiness under deployment-time attacks.
- Compared AutoGEO, MAGEO, AgenticGEO, and E-GEO to surface algorithmic assumptions and new optimization and defense directions.
- Proposed a three-component defense combining perplexity detection, Semantic Integrity Scoring, and Inter-Candidate Detection; achieved 1.000 precision and 0.000 false-positive rate across 1,200 Amazon ProductBench and MS MARCO evaluations.
- Developed six black-box attacks, showing that semantic manipulation via relevance inflation remains a structural blind spot for existing perplexity filters, safety classifiers, and paraphrasing defenses.
- Designed Tool Primitives and HEART's Planner-Router-Verifier workflow; built ToolFace, a repository of 25,519 functions supporting dynamic retrieval, nested calls, and feedback-driven recovery.
- Evaluated on five benchmarks with a 10% gain over SFT baselines and 6% over frontier commercial models, while reducing token consumption by up to 85%.
- Executed a 40,800-pair empirical study across 400 books and 17 prompt designs over six frontier models (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek V3, Qwen-Plus, Grok-4.1-Fast); uncovered a near-zero refusal rate (0.07%), revealing systematic cross-model safety-alignment gaps.
- Characterized the shift from refusal to warning-based moderation, with warning gaps of 8–15 points and prompt-dependent gaps reaching 19 points.
- Contributed to MonitorBench, a benchmark evaluating chain-of-thought monitorability across 1,514 instances, 19 tasks, and seven categories, with standard, direct-concealment, and monitor-aware-evasion settings.
- Implemented the impossible-coding-task evaluation adapted from ImpossibleBench and engineered a parallel multi-model inference and verification pipeline for large-scale experiments.
- Established semantic stability as a key reliability criterion for auto-generated prompts, with an optimization framework linking prompt consistency to system-level task success.
- Architected a full-stack multi-agent platform with editable and regenerable agent dialogues, reaching 95% task completion and reducing user intervention by 40%.
Research Systems & Platforms
- Developing GEO SaaS features for AI visibility auditing, content optimization, and SCI-Defense-based ranking-manipulation detection across ChatGPT, Gemini, and Perplexity.
- Engineering natural-language tool orchestration with dynamic routing, multi-turn clarification, browser-based integrations, and credential-isolated execution.
Professional Experience
- Resolved authentication and session-management defects across the Next.js frontend, Node.js/TypeScript backend, and Redis infrastructure by redesigning token-validation logic and decoupling protected routes from short-lived access tokens.
- Developed and deployed an IoT data-processing and machine-learning pipeline using Airflow, Pandas, TensorFlow, FastAPI, and C/Cython — reducing preprocessing time by 40%, improving model accuracy by 12%, lowering query latency by 30%, and accelerating inference by 10×.
Projects
- Developed a real-time AR inspection pipeline using Python, Open3D, and Unity — achieving 0.1–0.2 mm precision and ICP-based pose tracking for meshes with over 10 million vertices.
- Implemented an end-to-end PPO/GAE/KL-penalty pipeline with LoRA; achieved 70.1% pairwise preference accuracy and a 64.7% win rate versus 33.0% for the baseline on Stack Exchange data.
- Implemented Paxos and Raft over gRPC/Protobuf with 99.9% leader-election success and sub-200 ms election latency under network partitions; halved fault-detection time in simulated failures.
Teaching
- Tutored lower-year students in mathematics and computer science, supporting concept clarification, problem-solving strategies, and adaptation to a rigorous Technion-aligned curriculum.
- Graded assignments for Calculus, Topology, and Probability, giving feedback on mathematical reasoning, proof structure, and solution clarity.
Skills
LLM evaluation, benchmark design, adversarial testing, large-scale empirical analysis, RLHF, multi-agent system evaluation
Python, C/C++, Go, Java, TypeScript/JavaScript
PyTorch, vLLM, SGLang, LoRA, FastAPI, Docker, React, Node.js, gRPC/Protobuf