Machine Learning Engineer  /  Vision-Language-Action · World Models · RL

BorJiun Lin Christian

I teach machines to act from what they see. Then I go somewhere far away and photograph what they can’t.

Read the papers See the photographs

Education

National Tsing Hua UniversityBSc Computer Science · 2018 – 2022

Four years in Hsinchu. I started out in systems and drifted into machine learning after taking a semester of policy learning more or less by accident.

National Tsing Hua UniversityMSc Computer Science · 2022 – 2024

Thesis on sim-to-real transfer for manipulation. It became the CVPR workshop paper, and it is where I learned how to run an ablation honestly.

UCLAExchange Scholar · 2024

A term in Los Angeles working on video prediction for manipulation. Also the year I bought a real camera and started taking the photographs seriously.

Career

MicrosoftCloud Solution Architect Intern · 2022.08 – 2023.06
  • Integrated Azure Cognitive Services into proofs-of-concept for cost-effectiveness and efficiency against customer needs
  • Fine-tuned LLMs (OPT models) with the DeepSpeed framework on proprietary enterprise datasets, improving production efficiency for internal applications
  • Brought Azure OpenAI (GPT-3.5, GPT-4) into Microsoft Teams and ran fine-tuning on client-provided datasets to strengthen communication workflows
GoogleMachine Learning Engineer Intern · 2023.07 – 2023.10
  • Built a convolutional-recurrent model in C++ and Python for touchpad-gesture and mouse-movement recognition, deployed to internal tools with TensorRT for real-time response
  • Reached ~98% average classification accuracy at low memory usage
  • Ran the full pipeline through TensorFlow’s C++ API — dataset curation, model design, optimisation, deployment — portable to new scenarios for quick development
NVIDIAMachine Learning Engineer · 2025.03 – 2026.09
  • Fine-tuned VLMs (NVILA, Qwen2.5-VL, LLaVA) with GRPO on industrial video understanding — +54%, especially on instance action recognition
  • Few-shot video fine-tuning of the temporal-grounding model LITA on industrial datasets with RL policy optimisation (PPO, GRPO, GSPO), +35% IoU
  • Raised Cosmos-Reason 1.1 multi-action recognition accuracy to 97% via dynamic frame sampling and question augmentation for industrial SOP monitoring on limited data
RoboEraVLA / WAM Engineer · 2026.09 – present
  • Building vision-language-action and world-action models for humanoid platforms
  • Focused on long-horizon manipulation and cross-embodiment transfer
Bor Jiun Lin

Fig. 01 — somewhere with bad wifi

01 — About

I’m an ML engineer working on vision-language-action models and learned world models — systems that build an internal picture of a scene and use it to decide what to do next. Most of my day is spent on reinforcement learning in simulation, and the long argument about whether any of it survives contact with a real robot.

Before this I trained as an engineer with no intention of touching robots. A single semester of policy learning changed that. I’ve been chasing the gap between prediction and action ever since.

The rest of the time I’m outside. I photograph landscapes and the slow things in them — weather, light, ice, moss. It’s the same interest, really: watching a system change state and trying to be there at the right moment.

03 — Recent frames

Torres del Paine, CL
Mar 2026
Lofoten, NO
Feb 2026
Taroko Gorge, TW
Oct 2025
Namib Desert, NA
Aug 2025
Full gallery

04 — From the diary

12 Jul 2026 · 9 min

The benchmark is not the task

28 May 2026 · 6 min

What robots taught me about patience

All entries