Qwen3.5-VL (2B) — OCT/FPI action model
The first-generation vision-language-action policy for Meca500 trocar needle insertion into an eye phantom. A Qwen3.5-VL 2B backbone is extended with a sensor encoder that ingests two optical signals from the needle tip:
- OCT (optical coherence tomography) — A-line depth
- FPI (Fabry-Pérot interferometer) — contact force, read out of the interference phase
Those fuse with the visual and linguistic representations so the policy acts on what the tip is touching, not only on what the camera sees — which matters because at insertion depth the tip is no longer visible.
Contents. A single checkpoint, qwen_vla_final_1000.pt. There is no inference
wrapper, config, or tokenizer here; this is a training artifact kept as a record of the
Qwen generation.
Superseded
This line was replaced rather than extended. The successor rebuilt the stack on LeRobot (SmolVLA · ACT · Diffusion Policy · π0) with a MuJoCo digital twin, sharing almost no code with it. Both generations are kept because each holds work the other never absorbed.
Related public work: the dataset
Najongs/meca500-needle-insertion
(1,281 episodes, 145 GB) comes from the same robot and task. Research code for both
generations is private until publication.
Note on the name. This repository was previously called
Qwen2.5-VL_3B_OCT_FPI_Action_Model, which had both the model family and the parameter count wrong. The backbone is Qwen3.5-VL at 2B.
Jongyeol Na — Researcher, KIRO · M.S. DGIST, IROM Lab GitHub · Knowledge graph · W&B