My work includes fine-tuning robot policies, building multimodal robot datasets, and evaluating machine-learning systems. I’m looking for research engineering opportunities in robotics.
I’m especially interested in robot learning, multimodal data, vision-language-action models, simulation, and evaluation.
EXPERIENCE & EDUCATIONLuma AIPhysical AI Data CoAllen Institute for AIMicrosoftUMass Amherst
01 / ROBOTICS
Selected robotics work
Robot learning, multimodal data infrastructure, and simulation.
FEATURED / DATA INFRASTRUCTUREFeb–June 2026
Multimodal robot data platform
Co-founded Physical AI Data Co and built a collection, alignment, and replay stack for industrial robot arms.
Live session replay of a lentils grasp: RGB, depth, overview, DIGIT tactile, and aligned force/torque — open the dashboard ↗
My contribution. I designed the time-alignment pipeline that turns asynchronous RGB, tactile, force/torque, joint-state, and action streams into coherent trajectories, and built the visualization / replay path used to inspect missing or delayed samples before training.
Example replay (lentils): 1,899 frames / 63.3 s at 30 fps with 2,578 wrench samples. Streams that record at different rates are timestamp-aligned; delayed or dropped samples show up in replay before a trajectory is marked training-ready.
Output
Post-processed, labeled trajectories (e.g. sand, lentils, wool grasps) consumed as training data rather than raw logs.
Fine-tuned MolmoAct2 on teleoperated demonstrations and deployed the learned policy on a real SO-101 arm to strike xylophone bars.
Teleoperated demonstration from my LeRobot dataset (top camera, 30 fps) — human-operated, not a policy rollout.
My contribution. I collected LeRobot demonstrations (top + side cameras, 30 fps), ran the 8k-step MolmoAct2 imitation-learning fine-tunes, implemented the inference sequencer that maps a song to the eight trained instruction strings, and visualized rollouts in Rerun.
Experiment summary
Task
Single-note xylophone strikes on SO-101, sequenced into songs at inference time
Team award: First prize, Rerun track · Mission Robotics Hackathon
EMBODIED AI / SIMULATIONICLR 2026
Virtual Community
An open-world multi-agent simulator for humans, robots, and society — accepted at ICLR 2026.
The platform combines a physics-based multi-agent simulator with real-world-aligned 3D scenes. It supports two research challenges: community-scale planning among agents, and multi-robot cooperation with heterogeneous embodiments (humanoids, quadrupeds, drones, mobile bases).
My contribution. I was a co-author on the UMass / MIT-IBM team. I contributed to the embodied simulation platform and to the speed and memory benchmarks across agents, scenes, and resolutions — not the full system.
Problem: comparing generative image models across prompts and parameters requires running each variation manually
Approach: parameter sweep tool — mark any input field as swept, run variations in parallel, compare results in a labeled grid; Claude Sonnet expands prompts into stylistic variations along a named axis
Stack: Python • FastAPI • HTMX • Claude Sonnet • Replicate API • SQLite • Docker • Fly.io
Result: deployed app supporting 6 image models across 5 vendors with dynamic schema-driven forms
Co-developed a declarative probabilistic framework for evaluating semantic constraints over language-model token predictions using weighted model counting, calibrated token-event probabilities, and probabilistic consistency metrics.
Open-world multi-agent simulator for humans, robots, and society, with planning and multi-robot cooperation challenges. I contributed to the platform and to speed/memory benchmarks across agents, scenes, and resolutions. Accepted at ICLR 2026
Comprehensive analysis of pronoun and occupational biases in LLMs using statistical tests. GPT-4o: <5% non-preferred pronoun selection; Qwen2.5: 100% positional bias. Submitted to COLM'25
03 / ABOUT
About & interests
I have fine-tuned robot policies, built multimodal robot datasets, and evaluated machine-learning systems in production. I am looking for research engineering roles in robotics. I bring four years of engineering experience at Microsoft and an MS in Computer Science from UMass Amherst.
01
Robot learning
VLA fine-tuning, imitation learning, and real-robot deployment.
02
Multimodal data
RGB, tactile, force, joint-state, and action data aligned for training.
03
Research interests
World models, simulation, and evaluating how learned policies generalize.
04 / EXPERIENCE
Experience
Member of Technical Staff | Luma AI | July 2026 - Present
Developing generative-media systems across image-layering agents, inference pipelines, evaluation, and production infrastructure.
Co-Founder | Physical AI Data Co | Feb 2026 - June 2026
Built an end-to-end multimodal data platform for UR-series and xArm robots across RGB, tactile, force/torque, joint-state, and action signals.
Timestamp-aligned asynchronous streams and replayed trajectories so delayed or missing samples could be filtered before training.
Research Extern | Allen Institute for AI (AI2) | Jan 2026 - May 2026
Co-developed a declarative probabilistic framework for evaluating semantic constraints over language-model token predictions using weighted model counting.
Co-authored a paper accepted at EMNLP 2026 on calibrated token-event probabilities and probabilistic consistency metrics.
Machine Learning Engineer Intern | System1 | May 2025 – Aug 2025
Built a multi-agent Text-to-SQL system over enterprise data using RAG and execution-based feedback validation, achieving 98% query execution success; containerized experiments with Docker for reproducible iteration.
Designed automated evaluation and regression pipelines for model and prompt changes.
Software Engineer | Microsoft, Bing Ads | June 2020 – July 2024
Improved large-scale offline ad simulation and ranking evaluation over 2TB/day production data, reducing compute cost 9% and runtime 15% through sampling and evaluation optimizations.
Built a regression testing framework for parallel feature branch evaluation, reducing manual validation effort by ~90% and enabling systematic model and code comparisons.
98%
NL→SQL execution successSystem1 internship
2TB
Daily production dataMicrosoft · Bing Ads
9%
Reduction in compute costMicrosoft · Bing Ads
90%
Less manual validation effortMicrosoft · Bing Ads
05 / COMMUNITY
Recognition & community
Panelist — Building Data Agents Enterprises Can Trust
Open Future Forum · Palo Alto — Shared my experience building data agents and discussed enterprise AI on a panel with Jerry Xu.