Algorithm Engineer @ ByteDance

Zexi Cui

LLM Application Algorithm Engineer focusing on general-purpose Agent systems. Exploring the frontier of LLM, Agent, and Multimodal technologies.

LLM Agent Multimodal RAG
Recent readings

Background

Algorithm Engineer at ByteDance, working on LLM application and general-purpose Agent systems. My research interests span Large Language Models, Agent architectures, and Multimodal AI. I'm passionate about building intelligent systems that can reason, plan, and act autonomously.

2025.2 — Present
ByteDance — LLM Application Algorithm Engineer
Focusing on general-purpose Agent system design and development.
2023 — 2025
University of Sydney — Master of Data Science
Majoring in LLM, Agent, and Multimodal research.
2019 — 2023
Shanghai University — B.S. in Information Systems
Built foundation in computer science, data structures, and information systems.
Zexi Cui

Recent Papers

Papers I've been reading recently. Continuously updated.

ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, et al. — ICLR 2023
Agent Reasoning Tool Use
arXiv 2210.03629
Mar 2026
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, et al. — NeurIPS 2023
LLM Tool Learning
arXiv 2302.04761
Mar 2026
A Survey on Multimodal Large Language Models
Shukang Yin, Chaoyou Fu, Sirui Zhao, et al. — arXiv 2024
Multimodal Survey MLLM
arXiv 2306.13549
Feb 2026
Voyager: An Open-Ended Embodied Agent with Large Language Models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, et al. — NeurIPS 2023
Agent Embodied AI LLM
arXiv 2305.16291
Feb 2026
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, et al. — NeurIPS 2022
LLM Prompting Reasoning
arXiv 2201.11903
Jan 2026

Notes & Writings

Technical notes, project reflections, and thoughts on AI.

2026.03.01
Building a General-Purpose Agent: Lessons from Production

Reflections on designing and deploying Agent systems at scale. Covering architecture decisions, failure modes, and what actually works.

Agent Engineering
Read post
2026.02.15
RAG vs. Long Context: When to Use What

A practical comparison of retrieval-augmented generation and long-context models. Benchmarks, trade-offs, and production considerations.

RAG LLM
Read post
2026.01.20
Multimodal Understanding: From Vision to Action

How multimodal models are evolving beyond simple captioning into genuine understanding and grounded action in the real world.

Multimodal Research
Read post
2025.12.10
My Journey: From Information Systems to LLM Engineering

How I transitioned from studying information systems at SHU to working on cutting-edge AI at ByteDance, via data science at USYD.

Career Personal
Read post

Tech stack

AI / ML

LLM Agent Systems RAG Multimodal Prompt Engineering Fine-tuning RLHF LangChain LlamaIndex

Languages & Frameworks

Python PyTorch Transformers FastAPI TypeScript SQL

Tools

Git Docker Linux Wandb vLLM

Get in touch

Interested in collaboration, research discussion, or just want to connect? Feel free to reach out.

clay_cui@163.com