Qwen3.5-VL (2B) — OCT/FPI action model

The first-generation vision-language-action policy for Meca500 trocar needle insertion into an eye phantom. A Qwen3.5-VL 2B backbone is extended with a sensor encoder that ingests two optical signals from the needle tip:

  • OCT (optical coherence tomography) — A-line depth
  • FPI (Fabry-Pérot interferometer) — contact force, read out of the interference phase

Those fuse with the visual and linguistic representations so the policy acts on what the tip is touching, not only on what the camera sees — which matters because at insertion depth the tip is no longer visible.

Contents. A single checkpoint, qwen_vla_final_1000.pt. There is no inference wrapper, config, or tokenizer here; this is a training artifact kept as a record of the Qwen generation.

Superseded

This line was replaced rather than extended. The successor rebuilt the stack on LeRobot (SmolVLA · ACT · Diffusion Policy · π0) with a MuJoCo digital twin, sharing almost no code with it. Both generations are kept because each holds work the other never absorbed.

Related public work: the dataset Najongs/meca500-needle-insertion (1,281 episodes, 145 GB) comes from the same robot and task. Research code for both generations is private until publication.

Note on the name. This repository was previously called Qwen2.5-VL_3B_OCT_FPI_Action_Model, which had both the model family and the parameter count wrong. The backbone is Qwen3.5-VL at 2B.


Jongyeol Na — Researcher, KIRO · M.S. DGIST, IROM Lab GitHub · Knowledge graph · W&B

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading