feat(internvla-n1-dualvln): add TensorRT-Edge-LLM quantize and run instructions (PR #193) - #2
Merged
Merged
Conversation
…atrix Adds recipes/internvla-n1-dualvln/, the front matter for porting a working FP8 quantization and TensorRT-Edge-LLM deployment pipeline for InternVLA-N1-DualVLN on Jetson Thor. This commit is structure and documentation only; the Python and shell entrypoints land in follow-ups. The recipe is laid out along its dependency boundary rather than by file type. InternVLA-N1-DualVLN declares model_type internvla_n1 and ships no modeling code, so it cannot be loaded with trust_remote_code -- the class has to come from the InternNav repository. A repackaging step strips the eight System-1 tensor prefixes and rewrites config.json to a stock Qwen2.5-VL, using nothing but safetensors surgery, which takes InternNav off the hot path for quantization, export, engine build and latent verification. Only System-1 export and the agent-level checks need INTERNNAV_PATH, and the layout says so per step. configs/schemes.yaml encodes the scheme x strategy validity matrix so impossible combinations are rejected before the model loads rather than forty minutes into a run. Two entries are hardware facts: NVFP4 cannot touch the vision tower because the ViT MLP intermediate_size is 3420 and 3420/16 is not an integer, and the KV cache stays FP8 under NVFP4 weights because NVFP4 KV needs sm100f while Thor is sm110. NVFP4 weights are marked experimental rather than supported: they quantize, export and generate fluent text, but z_latents cosine falls to 0.647, which breaks the System2 -> System1 bridge. For a navigation model the bridge is the acceptance metric, not text fluency, so the README leads with that distinction. The README also retracts an earlier claim carried in the source project that FP16 engines are broken on Thor and FP8 is therefore mandatory. The garbage came from a Myelin fc_h_fusion miscompile at TensorRT 10.13; with the workaround applied the unquantized FP16 engine is in fact the highest-fidelity variant measured. FP8 is recommended on size and latency, not correctness. numpy and scipy are deliberately absent from the pyproject extra. This recipe needs numpy 1.x while three other extras pin numpy==2.2.6, so rather than assert a resolution outcome that has not been tested here, they are installed as a documented post-sync step. Note the repository already carries three mutually exclusive transformers pins across extras, so this is not a new class of problem; run uv lock and check before concluding anything about it. Directory names avoid build/ and lib/, which the root .gitignore silently swallows. A recipe-level .gitignore covers the multi-gigabyte engine and ONNX artifacts, which nothing else in the repository excludes. (cherry picked from commit 7b62c50)
repackage_system2.py strips the eight System-1 tensor prefixes and rewrites config.json
to a stock Qwen2.5-VL. Verified on the real 16 GB InternVLA-N1-DualVLN checkpoint: 729
System-2 keys kept, 609 System-1 keys dropped, and the result loads through AutoConfig and
AutoProcessor as Qwen2_5_VLConfig / Qwen2_5_VLProcessor with no InternVLA fields left. That
is what takes InternNav off the hot path for everything except System-1 export.
Five bridge tensors (latent_queries, cond_projector.{0,2}.{weight,bias}) are written to a
separate bridge.safetensors instead of being dropped with the rest of System 1. They are
what the z_latents fidelity check consumes, so keeping them here lets quantize, export,
build and verify all run from the repackaged directory without reopening the 16 GB source.
quant_schemes.py loads configs/schemes.yaml and rejects impossible combinations before the
model is loaded, so a bad request costs seconds instead of dying mid-quantization. The
vision-tower check reads vision_config.intermediate_size from the checkpoint rather than
hardcoding 3420, so it stays correct for other Qwen2.5-VL sizes; on this checkpoint it
reports '3420 / 16 = 213.75' in the error. Verified against all four strategies and both
scheme families: fp8 passes everywhere, nvfp4 is blocked on s3/s4 and gated behind
--allow_experimental on s1/s2.
A free-space preflight refuses to start below 18 GB, and --free_source deletes each source
shard as its converted copy lands for machines that cannot hold both checkpoints.
(cherry picked from commit 0f828a7)
…l loader Completes the System-2 quantization path. Verified end to end on the real checkpoint: 16.6 GB repackaged System 2 -> 9.4 GB FP8, 1164 TensorQuantizers, 80s quantize plus 16s export, hf_quant_config reporting FP8 with lm_head and visual* excluded, and tokenizer.model_max_length still 8192 (calibration truncates at 512 as a call-time argument only; a leak there would silently cap inference sequence length). model_loader.py drops the InternVLA-N1 branch entirely rather than carrying it forward. Post-repackage the input is always a stock Qwen2.5-VL, so the branch, its two InternNav imports, the diffusers gradient-checkpointing monkeypatch and the unused _is_system2_checkpoint helper are all dead weight -- 70 lines removed, and with them the dependency that would otherwise sit on the critical path of every quantization run. quantize.py calls the validity gate before loading the model, so an impossible request costs seconds. Confirmed: nvfp4+s3 is rejected in under a second with the checkpoint's own '3420 / 16 = 213.75' and never touches the weights, while fp8+s1 loads in 14s and reports sm_110. --cfg is renamed --scheme, kebab-case flags become snake_case to match the other recipes, and the output tag no longer hardcodes qwen2.5-vl-7b. calibration.py gains an offline path. This machine sits behind a TLS-intercepting gateway, so load_dataset() fails certificate verification even though the parquet shards are already in the hub cache -- and HF's own offline mode does not help, because it still wants the dataset script. The loader now probes reachability and reads the cached shards directly, which is the difference between a working and a non-working calibration run on any air-gapped or proxied deployment machine, not just this one. (cherry picked from commit 260bf45)
… scales applied Accuracy benchmarking a quantized checkpoint in PyTorch needs it loaded correctly, and the obvious way is silently wrong. export_hf_checkpoint writes LLM weights as real torch.float8_e4m3fn with dequantization scales in separate weight_scale/input_scale tensors. AutoModelForImageTextToText.from_pretrained on that directory appears to succeed -- it warns the scale tensors 'were not used when initializing' and continues -- but every quantized weight then sits at its raw FP8 magnitude. Measured on this checkpoint, the naive load produces weights 630x to 1799x too large depending on the layer. Any accuracy number taken from it would be reporting a broken load as quantization damage. load_quantized.py reconstructs w_bf16 = w_fp8.to(bf16) * weight_scale over a normal from_pretrained (which is what materialises the rope buffers; a meta-device build leaves them empty). Verified against the unquantized System 2 in float64: cosine 0.99965 and 2.6-2.7% relative error per projection, which is what FP8 E4M3 per-tensor should cost, against a scale ratio of exactly 1.000. Modules excluded from quantization -- the vision tower and lm_head -- are stored bf16 and pass through untouched. Note what this does and does not model: weight quantization only. Real FP8 W8A8 also quantizes activations, which the TensorRT engine does and this does not, so a number from here is a lower bound on the engine's deviation rather than a prediction of it. The docstring says so, because the distinction is easy to lose. (cherry picked from commit 63f376e)
… episodes
The official metrics (SR, SPL, NE, OS, nDTW) are closed-loop and need Habitat or
InternUtopia plus MP3D scenes. Neither is installed here and neither is practical on a
Jetson, so this measures what quantization actually threatens and what can be measured
offline: whether System 2 still picks the same waypoint.
Result on 42 held-out samples from two scenes, BF16 System 2 versus its FP8 quantization,
identical samples and greedy decoding:
pixel_goal_l2 BF16 mean 47.24 px / median 27.05 px
FP8 mean 50.12 px / median 26.97 px
agreement 31/42 replies byte-identical, median deviation 0.00 px
The mean moves 2.88 px while the median does not move at all, because the gap is carried
by 8 samples out of 42 rather than by a systematic shift -- worst case 89 px. That
distinction matters for a navigation model and is why both statistics are reported.
Two harness bugs were found by disbelieving the first numbers rather than publishing
them, and both are now prevented in code:
Feeding a pitched frame through a turn-1 prompt gave a 0% coordinate parse rate. The
model was right and the harness was wrong: the agent runs two turns, and the coordinate
only exists in turn 2, after the level view has been answered with a look-down token and
the tilted frame appended. The pixel goal in the level camera is the [-1,-1] sentinel on
every frame precisely because a level camera cannot see a point on the floor.
prompt_builder's 'episode_idx' is the step index within an episode, not the episode
number -- it derives the history frames from linspace(0, episode_idx-1). Passing the
episode number produced one image placeholder against nine images.
action_accuracy is now scored on turn 1 only. At turn 2 the model emits a coordinate by
design, so scoring it against a movement action reads 0% and looks like a regression.
(cherry picked from commit 08a64f6)
…port/build script README now leads the results with the number a navigation model is actually judged on: FP8 keeps the median waypoint error unchanged (27.05 -> 26.97 px over 42 held-out samples) with 31 of 42 replies byte-identical and a median deviation of 0.00 px. The 2.88 px gap in the mean comes from 8 disagreeing samples rather than a systematic shift, so both statistics are given; reporting the mean alone would overstate the damage and the median alone would hide the tail. It also states plainly what the measurement does not cover -- activations are quantized in the engine but not on this PyTorch path, so the figure is a lower bound -- and that the published SR/SPL/NE table is a literature reference, not something reproduced here, since the closed-loop metrics need Habitat or InternUtopia and MP3D scenes. 03_export_build_system2.sh collapses the source project's four near-duplicate export scripts into one, driven by --no_quantization for the FP16 reference. Two things it does are load-bearing and carry comments saying so: emit_hidden_states=True before export, without which the engine returns logits only and the System 1 bridge cannot be evaluated; and __LUNOWUD=-peep:fc_h_fusion=off on the build, without which an FP16 engine on TensorRT 10.13 produces fluent gibberish from a Myelin miscompile. The visual encoder is built once at 4/4096/1024 image tokens. The source project's default of 512 came from a single-image demo and cannot hold a VLN prompt's 9-10 frames, which is why its multi-image verifications silently pointed at a hand-built engine no script produced. (cherry picked from commit 3d4e764)
…onfig.json Two build-script defects found by running it rather than reading it. visual_build appends its own 'visual' component directory under --engineDir, so passing $ENGINE_DIR/visual produced $ENGINE_DIR/visual/visual/visual.engine. Every consumer looks for $ENGINE_DIR/visual/visual.engine, so the engine was built correctly and then filed where nothing would find it. Pass the parent instead. llm_bench reads its engine configuration from base_config.json while llm_build writes config.json -- same content, different name. Without the copy the benchmark refuses to open the engine that was just handed to it, with a parse error that reads like engine corruption rather than a missing filename. Verified on the FP8 engine built from the repackaged System 2: llm.engine 7.62 GB (the figure the source project recorded), visual.engine 1.36 GB, text smoke test coherent, and llm_bench measuring 95.8 ms prefill at 1024 tokens and 37.35 ms decode at 1024 past-KV. (cherry picked from commit 0eb6e91)
…ison
Both engines built from the same repackaged System 2 and measured with llm_bench on
Jetson Thor at batch 1:
LLM engine prefill 1024 decode pastKV 1024
base FP16 15.0 GB 196.22 ms 86.16 ms
FP8 7.62 GB 95.80 ms 37.35 ms
gain 1.97x smaller 2.05x faster 2.31x faster
Both emit the same text on the same prompt, and the PyTorch-side accuracy run already
showed the median waypoint error unchanged, so FP8 here is a straight win rather than a
trade: half the engine, roughly double the throughput, same decision.
Worth recording that the FP16 engine is correct only because the build applies
__LUNOWUD=-peep:fc_h_fusion=off; the build log confirms both workaround flags were passed.
Without it TensorRT 10.13 miscompiles Myelin's horizontal gate/up fusion on sm_110 and the
engine produces fluent gibberish, which is what once made FP8 look mandatory on this
platform. It is not -- FP8 is chosen here on size and speed.
(cherry picked from commit 997cdf6)
… both engines verify_latents.py rebuilds the bridge end to end against a PyTorch BF16 reference -- embed, scatter image embeddings, append the four trajectory tokens, run the engine with hand-built 3D mRoPE, take the last-layer hidden states, apply the host-side norm and cond_projector -- and compares z_latents. engine hidden pre-norm post-norm z_latents rel-L2 base FP16 0.999843 0.999123 0.999471 0.0293 FP8 (s1) 0.997793 0.987343 0.991861 0.1023 Both clear the 0.99 gate, so FP8 preserves the signal System 1 actually consumes. This is the check NVFP4 fails at 0.647 while still producing fluent text, which is why it ships as experimental rather than supported. The check reads latent_queries and cond_projector from the bridge.safetensors that repackage_system2.py sets aside, so it needs neither the 16 GB original checkpoint nor InternNav. Four transformers 5.x incompatibilities were fixed with version-probing helpers rather than a version pin, since the recipe should survive either: the vision tower and get_rope_index moved onto the inner model; get_image_features returns a BaseModelOutputWithPooling whose pooler_output holds the merged embeddings while last_hidden_state is the pre-merger tensor at vision width; and get_rope_index gained a required mm_token_type_ids argument. That last one does not fail on arity -- it fails later inside with 'NoneType is not an iterator', which is a poor way to discover it. Also switched the reference forward from flash_attention_2 to sdpa, which is what is available on Jetson and what the deployed agent uses. (cherry picked from commit b9b063f)
…nchmarks Adds traj_dit and memory-block export, their BF16 engine builds, the System-1 fidelity check and the two System-1 benchmarks, plus internvla_compat.py which carries the three patches needed to load System 1 at all: the DepthAnythingV2 checkpoint path, the diffusers gradient-checkpointing signature, and the traj_dit FFN multiplier. That last one is a real checkpoint-compatibility finding -- DualVLN was trained with ffn_dim_multiplier 2/3, and the stock build_traj_dit never passes it, so the state dict mismatches without the patch. Four defects carried over from the source project are fixed here rather than propagated: MemBlock was defined inline in the memory exporter as well as in memblock.py, so the two copies could drift -- the exporter now imports the single definition; the traj_dit engine filename said _async in one place and not the other, and neither constant was read; a dead 'if False else None' statement; and INTERNNAV_ROOT is renamed to the recipe-wide INTERNNAV_PATH. trt_torch.py keeps its original NVIDIA Apache-2.0 header rather than being restamped. It is third-party code and Apache-2.0 section 4 requires retaining the notice; it is the one file in this recipe that does not carry the VinRobotics BSD header. Documented an environment constraint found by hitting it: System 1 must run under transformers 4.51.3. The InternNav modeling code reads config.hidden_size off the top-level config, which transformers 5.x no longer flattens, so it fails with a bare AttributeError. The System-2 path is unaffected and runs on either version. (cherry picked from commit 4707c05)
Both System-1 engines built and verified loadable with the expected I/O: traj_dit 134 MB ONNX -> 72 MB engine x, timestep, z_latents -> output memory block 200 MB ONNX -> 104 MB engine images -> memory_tokens Sizes match what the source project recorded. Both stay BF16 on purpose: they are small enough that quantizing them buys nothing, and the diffusion head is the part least tolerant of it. The memory exporter no longer exits 1 after a successful build. Its in-script parity check imports the TensorRT Python bindings, which JetPack ships for Python 3.12 only, while the export itself needs transformers 4.51 from the 3.10 environment -- so the check could never run in the same interpreter as the export. It now skips with a message pointing at verify/verify_system1.py under the 3.12 environment. Reporting failure for an artifact that was produced correctly is worse than not running the check. Two more environment gaps closed by installing rather than working around: diffusers 0.33.1 (the traj_dit scheduler) and setuptools<81 (InternNav's LongCLIP still imports pkg_resources, which newer setuptools dropped). (cherry picked from commit 8722391)
…hmarks Adds the five remaining fidelity checks, the System-2 benchmark, the simulator adapter and its eval entry point. verify_engine_policy.py runs green on the FP8 engine: the adapter is a proper InternVLAN1Net subclass adding only _engine_generate, and the engine's reply survives the tokenizer roundtrip exactly. lib/_trt_contract.py is deliberately not ported. It is dead code -- imported by nothing, reading an export_metadata.json that no script writes, and referencing a different project's exporter. Defects fixed rather than carried across: two scripts still defaulted to one developer's absolute layout under vln-thor/artifacts; benchmark_system2.py wrote reports/bench_e2e.json without creating the directory, so it crashed on its own last line after doing all the work; INTERNNAV_ROOT/VLN_OPT_WORK/VLN_OPT_ENGINES are renamed to the recipe-wide INTERNNAV_PATH/WORK_DIR/ENGINE_DIR; nine Vietnamese comments translated; and seven files restamped to the BSD header while trt_torch.py keeps its NVIDIA Apache-2.0 notice. benchmark_system2.py depends on a golden manifest that neither this recipe nor the source project generates. Instead of leaving that as a FileNotFoundError on the first read, it now fails up front and describes the file it wants -- the same sample shape quantize/benchmark_accuracy.py already builds from LeRobot episodes. Scripts that need prompt_builder now locate the recipe root by walking up to the directory containing quantize/, rather than counting parent levels: they sit at two different depths and a fixed count works for one and not the other. (cherry picked from commit b24d213)
… not the weights
Measured in PyTorch on a real NVFP4 checkpoint of this model, weights only:
weight rel-err weight cos z_latents
FP8 2.67% 0.999644 0.998020
NVFP4 9.45% 0.995534 0.987986
NVFP4 weight quantization costs 0.988, not 0.647. The weight error is 3.5x FP8's and the
bridge degrades roughly in proportion, which is unremarkable. So the 0.647 recorded for the
NVFP4 engine does not come from weight quantization, and the candidates are the two things
this measurement does not model: 4-bit activation quantization, and the engine itself.
The engine hypothesis is the one I would chase first. This platform has already produced
one 'quantization is broken' conclusion that turned out to be a TensorRT miscompile, and
there is a second known one specific to NVFP4 -- CASK miscompiling two or more fused
epilogues in an NVFP4 GEMM at batch 1, fixed by -cask_fusion:max_num_epilogues=1. A 0.647
taken from an engine built without that workaround would be measuring the miscompile rather
than the format.
Channel analysis rules out the obvious remedy for the weight-side loss: at the final layer
the top 128 channels by magnitude carry only 29.5% of the squared error, and masking them
lowers cosine rather than restoring it. The error is spread, not concentrated in outliers,
so AWQ scaling or a targeted exclusion has nothing to grip.
load_quantized.py gained NVFP4 support to make this measurable at all: unpacking two E2M1
values per byte with a per-16 FP8 block scale and a float32 global scale. The quantized path
now hands transformers an already-dequantized state_dict rather than letting it read the
checkpoint -- FP8 loads silently wrong without scales, and NVFP4 fails outright on the
halved width, so from_pretrained cannot be trusted with either.
NVFP4 stays experimental. The next step is to rebuild its engine with the CASK workaround
and re-measure end to end, not to do more weight analysis.
(cherry picked from commit 46ff847)
…t standard
The recipe carried 862 flake8 warnings while every existing recipe in this repository is
clean at zero. That gap is now closed. Most of it was VLN-Opt's dense one-liner style --
compound semicolon statements, missing whitespace, multi-import lines -- mechanically
fixable with autopep8 at the repo's 120-column setting.
Three defects were real, not cosmetic, and only surfaced because the lint pass ran:
* benchmark_system2.py did not parse at all. A sys.path shim had been inserted directly
above an import that lives inside a function, so its module-level indentation broke
the enclosing block. The file had never been executed after that edit.
* verify_latents_vln.py referenced json without importing it -- left behind when
load_ckpt_tensor was rewritten to read bridge.safetensors, since the old body carried
the import.
* Six unused imports and one dead local.
E402 is silenced with explicit noqa rather than reordered: these modules must insert a
sys.path entry before importing prompt_builder or the InternNav tree, so the import
genuinely cannot come first.
Verified the reformatting changed no behaviour: the scheme validity gate still accepts
fp8+s1, rejects nvfp4+s3 with the checkpoint's own 3420/16 arithmetic, and gates nvfp4+s1
behind --allow_experimental; the dequantizing loader still reproduces cosine 0.999647 for
FP8 and 0.995489 for NVFP4 against the unquantized weights.
Separately confirmed the NVFP4 nibble order empirically rather than trusting the spec: the
low-nibble-first unpacking gives cosine 0.995485 against the reference weights while the
swapped order gives -0.002446, so the layout is not a coin flip that happened to land.
(cherry picked from commit 0679e0b)
…t of its collapse Rebuilding the NVFP4 engine with the fork's CASK workaround settles most of the question the investigation opened. Three measurements, each isolating a layer: FP8, weights only (PyTorch) 0.998020 NVFP4, weights only (PyTorch) 0.987986 NVFP4 engine, with the CASK workaround 0.931005 NVFP4 engine, no workaround (source figure) 0.647 So the 0.647 was mostly a compiler artifact. -cask_fusion:max_num_epilogues=1 -- applied automatically by the fork, gated to NVFP4 graphs at batch 1 -- moves it to 0.931, and the build log confirms all three flags fired. Weight quantization was never the problem: NVFP4 weights cost 0.988, an error 3.5x FP8's with the bridge degrading in proportion. A real gap remains between 0.988 and 0.931, and it now has a named suspect rather than being a mystery: NVFP4 is W4A4, and 4-bit activations through a 3584-wide hidden state in blocks of 16 are the part the PyTorch weights-only path cannot model. The 0.647 is quoted from the source project and was not reproduced here. What is measured here is that the same checkpoint reaches 0.931 once the workaround is applied. NVFP4 stays experimental: 0.931 is below the 0.99 gate. But it is far from broken, and the 5.5 GB engine is the smallest of the three. Adds 06_measure.sh to collect prefill/decode latency for every built engine into the JSON run_matrix.py reads, and run_matrix.py itself, which assembles sizes, latency, bridge fidelity and task accuracy into one table without recomputing anything. (cherry picked from commit bd71e87)
All three variants built from one repackaged System 2 and measured on an idle GPU: variant ckpt LLM engine prefill decode z_latents pixel L2 mean/median BF16 16.6 GB 14.15 GB 135.8ms 56.4ms 0.999471 47.24 / 27.05 px FP8 s1 10.1 GB 7.62 GB 82.1ms 31.5ms 0.991861 46.26 / 22.51 px NVFP4 s1 7.2 GB 4.77 GB 73.2ms 20.2ms 0.931005 40.69 / 23.54 px FP8 is the recommendation: 1.86x smaller and 1.65x/1.79x faster than BF16, bridge held at 0.9919, and the median waypoint error does not worsen. It reads slightly better (27.05 -> 22.51 px), which for 42 samples should be read as unchanged rather than as a gain from quantization. NVFP4 is smaller and faster still but its bridge sits at 0.931, below the gate, so it stays experimental. Two caveats are documented alongside the table because both cost time to learn here. Latency must be measured on an idle GPU: sharing the device with another job inflated FP8 prefill from 82 ms to 117 ms, and nothing in the scripts enforces this. The 'identical replies' count is a weak indicator and is now labelled as such. It moved from 31/42 to 11/42 for FP8 between two revisions of the checkpoint loader while pixel_goal_l2 barely moved. I initially attributed that to GPU contention; a clean re-run reproduced the contended numbers exactly, so the cause was the loader change, not contention. The current loader is verified correct -- all 729 weights in the loaded model match the dequantized state dict -- so the current numbers stand. Greedy decoding over a 152k vocabulary flips on tiny logit differences; the L2 median is the number to trust. (cherry picked from commit b60e7d3)
AWQ does not help this model. End to end the bridge measures 0.986293 against plain NVFP4's 0.987986 -- the same within noise. That matches what the channel analysis predicted: the error is spread across channels rather than carried by outliers, so AWQ's per-channel rescaling has nothing to grip. Getting to that number needed two loader fixes, and the first reading was wrong in a way worth recording. AWQ-lite scales activations by a per-input-channel factor and stores the weight pre-divided by it, so a plain matmul has to multiply it back in. Ignoring the pre_quant_scale tensors entirely produced ~190% relative weight error, which reads as AWQ being catastrophically broken rather than being loaded wrong. The direction is also easy to invert: verified empirically on this checkpoint that multiplying gives cosine 0.9898, leaving it alone 0.9606, and dividing 0.8089. The second is subtler and invalidates a metric rather than a number. Only down_proj and o_proj carry an explicit pre_quant_scale, because q/k/v share one layernorm and gate/up share another -- for those, AWQ folds the scale into the layernorm weight instead. Measured here, AWQ moves the layernorm weights by 0.012x to 91x while plain NVFP4 leaves them at exactly 1.000. So comparing AWQ's linear weights against the reference's linear weights is not a comparison at all; the invariant is the composition. The weights stage now detects AWQ checkpoints and skips with that explanation rather than reporting a meaningless 180%. qat.py adds quantization-aware fine-tuning, which is the remaining lever: the gap between NVFP4 weights-only (0.988) and the engine (0.931) is 4-bit activations, and no weight-side method reaches those after the fact. ModelOpt has no separate QAT entry point -- it is mtq.quantize followed by ordinary fine-tuning with the fake-quantizers left in so gradients pass through them. It excludes scene YmJkqBEsHnH from training by default: the same MP3D building appears as calib_scenes/r2r and probe_heldout/rxr, and the source project has already reported one calibration gain that turned out to be exactly this leak. --allow_overlap disables the guard for anyone who wants to measure the effect deliberately. Note that success rate cannot be measured on this machine -- SR, SPL and NE are all closed-loop. QAT is evaluated on the proxies it can move, z_latents and pixel-goal L2, and the docstring says so. (cherry picked from commit 332bf0b)
…weight-side gap
All three NVFP4 presets measure the same end to end against the unquantized reference:
nvfp4_default 0.987986 baseline
nvfp4_awq_full 0.986293 equal within noise
nvfp4_local_hessian 0.987986 byte-identical to default
nvfp4_local_hessian is a no-op on this model, and it is worth stating plainly because it
fails silently. The preset genuinely differs -- algorithm={'method': 'local_hessian',
'fp8_scale_sweep': True} against 'max' -- and the run exits 0, but the exported weights
match nvfp4_default bit for bit: 0 of 6,422,528 bytes differ, scales included. Two runs
with different algorithms cannot produce identical output unless the algorithm never ran.
It stays selectable via --scheme, so a user would reasonably believe they had tried it.
AWQ does run, its weights genuinely differ, and it still does not help -- which is what the
channel analysis predicted: the error is spread across channels rather than carried by
outliers, so per-channel rescaling has nothing to grip.
Three weight-side methods stopping at the same 0.988 is the argument for QAT rather than a
fourth. The gap from there to the engine's 0.931 is 4-bit activations, and no
post-training method reaches those.
(cherry picked from commit 1d1e8be)
…ne run
Adds quantization-aware fine-tuning and reports what it actually did, which was make
things worse:
z_latents (engine) pixel L2 mean / median
NVFP4 PTQ 0.931005 40.69 / 23.54 px
NVFP4 + QAT 0.891583 41.86 / 23.16 px
The training loss rose from 0.78 to 1.13 over 16 steps, so the run pushed the weights the
wrong way rather than converging. That makes this a failed training run, not evidence that
QAT cannot help: the experiment never reached the question it was meant to answer. Saying
'QAT does not work for this model' from sixteen steps would be a conclusion the data does
not support, so the README says what would change next instead -- a much smaller learning
rate, hundreds of steps rather than sixteen, and a warmup.
Two obstacles are worth recording because both cost a run to find.
Full fine-tuning does not fit on this hardware. Weights plus gradients plus AdamW moments
for 8.29 B parameters are ~100 GB before any activations, against a 122 GB pool shared with
the host; the first attempt was killed by the OOM killer. --train_last_n_layers (default 4)
makes it fit at 932 M trainable parameters, and it targets the right place anyway, since
the bridge reads the last layer's hidden states.
Freezing combined with gradient checkpointing needs two fixes applied together. The
activations entering the first trainable layer carry no grad_fn, and reentrant
checkpointing then discards the graph, so loss.backward() fails with 'element 0 of tensors
does not require grad'. Both use_reentrant=False and enable_input_require_grads() are
required; either alone still fails.
Also documents a metric trap I walked into first: z_latents against the unquantized
reference is the wrong measure for QAT. PTQ approximates the original model, so similarity
to it is meaningful; QAT deliberately moves away from the original to absorb quantization
noise, so a good QAT run can lower that number while improving behaviour. QAT is judged
through the engine and on task accuracy instead.
(cherry picked from commit 2f0b807)
…engine number
The results table carried two z_latents columns that measure different things, side by
side, with nothing saying so. That is a real reporting defect: it invites the reader to
treat 0.988 and 0.931 as disagreeing measurements of one quantity when they are correct
measurements of two.
load_for_eval reconstructs the weights into a plain model with no quantizers, so it
reproduces weight quantization error only -- activations stay bf16 and the matmuls are
bf16. But FP8 is W8A8 and NVFP4 is W4A4: the engine quantizes every activation too. The
weights-only path cannot see that half at all.
load_fake_quant now inserts live quantizers via mtq.quantize, so both halves are simulated
and a PyTorch number can be put next to an engine number honestly. load_for_eval stays for
weight-side questions, with its docstring saying which is which.
The README now shows the decomposition rather than the bare pair:
weights only engine activation cost
FP16 - 0.999471 -
FP8 0.998020 0.991861 0.006
NVFP4 0.987986 0.931005 0.057
The FP16 engine at 0.999471 is what makes this legible: TensorRT itself costs about 0.0005,
so almost none of the gap is export or runtime. Going from 8-bit activations to 4-bit costs
nine times more than going to 8-bit, which is the whole reason NVFP4 misses the gate here
while FP8 clears it.
(cherry picked from commit 9b87622)
…vation quantization Measuring fake quant directly refutes what this recipe previously claimed. With live quantizers in PyTorch -- weights and activations both simulated, as the engine does them -- NVFP4 measures 0.978631, against 0.987986 weights-only and 0.931005 through the engine. That splits the 0.057 total as: weight quantization 0.012, activation quantization 0.009, and everything else 0.048. Activations account for 16% of the gap. The remaining 84% only appears once the model runs as a TensorRT engine. I had attributed the whole gap to activations, reasoning from three numbers rather than measuring the middle one. The reasoning was plausible -- NVFP4 is W4A4, 4-bit activations through a 3584-wide hidden state are an obvious suspect -- and wrong. compare_fake_quant.py exists to make that check cheap, and its docstring states both outcomes up front so it cannot be run as a confirmation exercise. FP8 has no comparable effect: its entire PyTorch-to-engine gap is 0.006 against NVFP4's 0.048, and the FP16 engine at 0.999471 bounds TensorRT's generic cost at ~0.0005. So this is specific to the NVFP4 kernel path, the same neighbourhood as the CASK epilogue miscompile that -cask_fusion:max_num_epilogues=1 improves from 0.647 to 0.931 without evidently resolving. The practical consequence changes what to do next. NVFP4 quantization on this model is close to the gate at 0.9786; it is the engine that loses the rest. That is also why AWQ, local-Hessian and QAT all failed to move it -- all three attack quantization quality, and quantization quality was never the binding constraint. (cherry picked from commit f8557ef)
…ut batch size diagnose_engine_gap.py compares an engine against its own fake-quant model in one process on the same inputs, so the difference is the engine alone rather than engine plus quantization. Hidden states before the final norm, text prompt: FP8 0.998256 NVFP4, maxBatchSize 1 + CASK cap 0.986790 NVFP4, maxBatchSize 2, no CASK flag 0.986790 The NVFP4 engine carries roughly eight times FP8's engine-side error, 0.013 against 0.002, and the bridge amplifies it -- 0.013 in hidden states becomes the 0.048 seen in z_latents after the final norm, GELU and cond_projector. It is not the batch-1 miscompile. The two NVFP4 engines are genuinely different builds (different checksums, maxBatchSize 1 and 2, and the fork correctly withholds -cask_fusion:max_num_epilogues=1 from the batch-2 build) yet measure identically to six decimal places. So the CASK cap fully recovers what the batch-1 path loses, and the residual is inherent to the NVFP4 kernels. My guess that a second batch-1 miscompile was hiding here was wrong. The FP8 control earned its place. The first version of this diagnostic compared against output_hidden_states[-1] and read 0.4837 for NVFP4 -- which looks exactly like a broken kernel. FP8 read 0.4949 through the same path, and FP8 is known good at 0.9919, so the harness was wrong: the engine emits hidden states before the final norm while that tensor is after it. Hooking the norm and taking its input moves FP8 to 0.9983. Without the control I would have reported a broken NVFP4 kernel path for the second time in this investigation. (cherry picked from commit cbddbcc)
…sorRT
Checked upstream first, as asked: InternNav ships no TensorRT or ONNX export at all --
nothing under internnav/ references trtexec, tensorrt or torch.onnx. The System-1
conversion is entirely this recipe's, inherited from the source project.
Two engines carry the compute, and both were verified by execution rather than by
inspection:
traj_dit (NextDiT head) 134 MB ONNX -> 72 MB engine
x[64,32,384] f32, timestep[64] i64, z_latents[64,*,768] f32 -> output[64,32,384]
memory block (DAv2 + MemoryEncoder + QFormer) 200 MB -> 104 MB
images[T,3,224,224] f32 -> memory_tokens[1,32,768]
Finite output, correct shapes. Worth noting the engine I/O is fp32/int64 despite BF16
weights; feeding bf16 trips an assertion in the wrapper rather than converting.
action_encoder, action_decoder, pos_encoding, cond_projector and the flow-matching
scheduler loop stay in PyTorch on the host. That is a design choice, not an omission --
they are tiny or control-flow heavy.
The end-to-end parity check remains unrunnable here and the README now says why rather
than leaving it as 'ready'. verify_system1.py needs InternNav and TensorRT in one
interpreter. InternNav targets transformers 4.x: under 5.x it fails first on
config.hidden_size, which transformers now nests under text_config -- patched in
internvla_compat.patch_config_flattening -- and then on apply_chunking_to_forward, removed
from modeling_utils. That chain has no natural end, and downgrading transformers inside the
TensorRT environment risks the edgellm exporter that currently works. Running the check
needs a fourth environment with transformers 4.51 and the TensorRT bindings together.
(cherry picked from commit 2193c6b)
…n two stages The parity check needs InternNav (transformers 4.x, Python 3.10) and the TensorRT bindings (Python 3.12) at once, which no interpreter here has. Splitting it across the two environments removes the conflict: stage A dumps the PyTorch reference's inputs and outputs to a .pt, stage B feeds the engines those same tensors. Both engines reproduce PyTorch: memory_tokens 0.999981 traj_dit, one step 0.999508 full trajectory 0.999670 (rel-L2 0.0258) Stage A also captures the reference's starting noise and one real traj_dit call. Both matter: generate_traj draws its latents mid-function, so reseeding gives a different valid trajectory (cosine ~0.31), and the single-step probe is what separates a bad engine from a bad reimplementation of the sampler loop. MemBlock expects ResNet-normalized input because generate_traj normalizes before rgb_model; passing raw pixels made memory_tokens read 0.315 while each half stayed internally consistent.
…ss both systems The results section reported System 2 and System 1 separately, with no dtype column and no System-1 latency at all, so nothing said what a whole planning step costs. Adds System-1 latency (stage B now takes --bench_iters) and the PyTorch baseline to compare it against, then restructures the matrix into three tables: System 2 by quantization scheme, System 1 as a conversion result, and both together per planning step. Every row names its dtype -- weights, activations, KV cache and vision tower separately, since s1 leaves the last two alone. System 1 engines 63.1 ms vs 175.4 ms in PyTorch 2.78x full step, FP8 709 ms vs 1806 ms all-PyTorch 2.55x weights, FP8 9.16 GB vs 15.7 GB unquantized Latency is flat in num_sample_trajs on the PyTorch side (175.4 / 175.6 / 173.1 ms at 32 / 4 / 1), so System 1 is launch-bound, which is why engines pay there. Also drops the stale duplicate table and the superseded 12-sample figures, and fixes the stage B GELU to approximate="tanh" to match the checkpoint's cond_projector.
System 1 shipped BF16 on the assumption that quantizing it was not worth it. This measures that instead of assuming it, and the assumption holds -- for a reason that is worth writing down. Unlike System 2, System 1 goes torch.onnx.export -> trtexec, and TensorRT does FP8 only through explicit quantization, so the Q/DQ has to be in the ONNX. Adds a ModelOpt PTQ pass calibrated on real tensors captured from a live System 2 -> System 1 run rather than random draws: FP8 scales are amax-based and both modules consume tensors whose scale is set upstream, so a synthetic draw is quietly wrong. FP8 works -- 328/328 Q/DQ pairs in traj_dit, 160/160 in the memory block, 1.55x smaller, 20% faster on the diffusion loop -- and is still the wrong trade: engines 176 MB -> 114 MB 0.7% of 9.16 GB deployed step 710 ms -> 698 ms 1.7%, System 2 dominates waypoint dev 0.0032 -> 0.0198 6x, p95 39% of a waypoint's reach Splitting it shows where the loss comes from: traj_dit alone costs 0.9997 -> 0.9880 even though its per-step error is only 0.999508 -> 0.997811, because the sampler runs 10 steps and each feeds the next. The memory block adds the rest and buys nothing in latency (2.09 vs 2.04 ms). Adds waypoint deviation to the parity check -- cosine says how aligned two trajectories are, not how far apart the robot ends up -- plus --engine_suffix so the same check runs against either precision. Three things that fail quietly, all now handled: ModelOpt emits trt::TRT_FP8*, not ONNX QuantizeLinear, so a naive counter reports a quantized graph as unquantized; the legacy exporter cannot infer a conv kernel shape through Q/DQ; and the QFormer's fused MHA fast path has no ONNX symbolic.
… calls
Six of the seven scripts the Makefile invokes did not exist, so every entry point
failed at the first command -- including the README's own Quick Start. The Python
underneath was fine; only the shell wrappers were missing.
00_fetch_calib_scenes.sh pull a diverse InternData-N1 subset (gated dataset)
01_repackage.sh InternVLA -> stock Qwen2.5-VL System 2
02_quantize.sh one scheme x strategy, through the validity gate
04_export_system1.sh traj_dit + memory block -> BF16 engines
05_verify.sh both acceptance gates
benchmark retargeted at 06_measure.sh, which exists and takes
the engine parent directory via environment, not flags
05_verify.sh --system1 runs the two-stage split, so it takes PYTHON_PT and PYTHON_TRT:
InternNav wants transformers 4.x and the TensorRT bindings ship for Python 3.12 where
transformers is 5.x, and nothing reconciles the two in one interpreter.
Exercised rather than assumed. The System 2 gate reproduces the README's 0.991861 on
the FP8 engine; the System 1 gate reports 0.999661 and exits 0; a missing checkpoint
and nvfp4 x s3 both fail with their reason.
Also fixes two README flags that never existed (--calib_data, --max_seq_len; the real
one is --calib_data_root) and drops the unassigned-owner line.
…port NVIDIA/TensorRT-Edge-LLM#193 exports InternVLA-N1-DualVLN directly, no repackage step, bridge folded into the graph, plus an async C++ runtime and 199-episode closed-loop SR this recipe never measured. Also flags that the recipe's central claim -- z_latents cosine as the acceptance metric -- does not hold. Closed-loop SR showed neither z_latents cosine nor trajectory cosine predicts navigation success for this model, in either direction; they rank this recipe's own FP8 vs NVFP4 call the wrong way. The FP8-on-System-1 measurement and the NVFP4 root-cause analysis are unaffected and still hold.
hungho77
force-pushed
the
feat/internvla-n1-dualvln
branch
from
August 24, 2026 19:04
b9b8c46 to
e285a9a
Compare
…structions (PR #193) NVIDIA/TensorRT-Edge-LLM#193 (open, not yet merged) adds native InternVLA-N1-DualVLN export -- direct checkpoint export, the z_latents bridge (final_norm + cond_projector) folded into the graph -- so once building from that branch, the repackage pass, host-side bridge computation, and custom export/quantize scripts this recipe used to carry are no longer needed. What is left as this recipe's job is the one piece that stays outside TensorRT-Edge-LLM regardless: building a navigation-domain calibration set, since calibrating on the deployment prompt's own domain rather than generic news text measurably improves quantization quality (FP8 trajectory cosine 0.909 -> 0.978 in earlier testing). Replaces the old repackage/quantize/export/verify script tree (45 files) with a single build_calib_jsonl.py plus a two-target Makefile, and rewrites the README with quantize + export + build instructions for PR #193's branch, plus the 199-episode closed-loop SR results, which supersede the recipe's earlier z_latents-cosine-based FP8-over-NVFP4 recommendation.
hungho77
force-pushed
the
feat/internvla-n1-dualvln
branch
from
August 24, 2026 19:11
e285a9a to
19d9d5f
Compare
…int at PR #193 The file link resolved against TensorRT-Edge-LLM's main branch, where PR #193 hasn't landed yet, so it 404s. Link the PR itself instead of a path that only exists on its branch.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add
internvla-n1-dualvlnrecipe: quantize + run instructions for TensorRT-Edge-LLM (PR #193)Adds a recipe for InternVLA-N1-DualVLN, a dual-system vision-language navigation model,
targeting Jetson Thor (sm_110, JetPack 7.1, TensorRT 10.13.3.9).
Status
Native InternVLA-N1-DualVLN support is proposed upstream as
NVIDIA/TensorRT-Edge-LLM#193 — open,
not yet merged — tracked by
NVIDIA/TensorRT-Edge-LLM#190. Built
from that branch,
tensorrt-edgellm-quantize/tensorrt-edgellm-exportread the checkpointdirectly and fold the
z_latentsbridge (final_norm+cond_projector) into the graph — norepackage step, no host-side bridge computation. What this recipe adds is the one piece that
stays outside TensorRT-Edge-LLM regardless: instructions for quantizing and running the model
with TensorRT-Edge-LLM, plus a navigation-domain calibration set to feed into that flow.
Measured on Jetson Thor, 199 R2R val_unseen episodes, closed-loop:
Neither differs from PyTorch significantly — both FP8 and NVFP4 are viable; NVFP4 for size and
speed, FP8 for a wider margin. Full validation, including why
z_latentscosine and trajectorycosine were tried and rejected as acceptance criteria, is in PR #193.
What this recipe does
Until PR #193 merges, build TensorRT-Edge-LLM from that branch directly:
git clone https://github.com/NVIDIA/TensorRT-Edge-LLM.git cd TensorRT-Edge-LLM git fetch origin pull/193/head:internvla-n1 git checkout internvla-n1Then build a calibration set of realistic navigation prompts. Calibrating the quantized
backbone on its own prompt domain, rather than generic news text, measurably changes
activation scales:
huggingface-cli download InternRobotics/InternData-N1 \ vln_ce/raw_data/r2r/train/train.json.gz --repo-type dataset \ --local-dir $CALIB_DATA_ROOT python quantize/build_calib_jsonl.py \ --train_json $CALIB_DATA_ROOT/vln_ce/raw_data/r2r/train/train.json.gz \ --output $CALIB_DATA_ROOT/nav_calib.jsonlEverything after that is the standard TensorRT-Edge-LLM flow, in its own environment:
System 1 (the trajectory expert) and the async dual-system runtime
(
internvla_n1_dual_system_inference/internvla_n1_dual_system_server) are not part of thisrepo — they ship with TensorRT-Edge-LLM. See
NVIDIA/TensorRT-Edge-LLM#193 for the
full export → build → run flow.
Calibration text needs no trajectory-query placeholder tokens:
tensorrt-edgellm-quantizeloads System 2 as a stock Qwen2.5-VL (
internvla_n1_loader.py) and calibrates with ordinarytext forward passes, never touching the bridge — those tokens are registered later, at export
time, so appending them during calibration would only add noise.
What is and is not quantized
Quantized: the System 2 LLM backbone. Never quantized: System 1 (
traj_dit, memory block), andthe bridge (
cond_projector,latent_queries) — four rows through a Linear/GELU/Linear, keptat source precision because quantizing them saves nothing measurable and puts error directly on
the tensor System 1 steers by.
NVFP4 with the vision tower is blocked, permanently. The Qwen2.5-VL ViT MLP has
intermediate_size = 3420, and 3420 / 16 = 213.75 — not divisible by the NVFP4 block size. Onlythe LLM backbone is quantized here, so this does not apply, but it is worth knowing if this
recipe is ever extended to a strategy that includes the vision tower.
Tested environment
Jetson Thor, JetPack 7.1 (TensorRT 10.13.3.9, CUDA 13).
pip install -e ".[tools]"in theTensorRT-Edge-LLM checkout pulls
nvidia-modeloptanddatasets; without ittensorrt-edgellm-quantizefails at import.Note on this PR's history
Earlier commits on this branch built a full custom pipeline — repackaging the checkpoint into a
stock Qwen2.5-VL, computing the
z_latentsbridge on the host, and vendoring System 1 export —because at the time TensorRT-Edge-LLM had no InternVLA-N1 support at all. That work fed directly
into NVIDIA/TensorRT-Edge-LLM#193, which
proposes native support upstream and is still open. Until it merges, this recipe builds from
that PR's branch directly (see the git-fetch instructions above); this PR now ships only what
still belongs in a recipe repo regardless of when #193 lands: instructions for the
quantize/export/build flow, and the navigation-domain calibration set.