Ship the quantized kernels as their own library - #21642
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21642
Note: Links to docs will display an error until the docs builds have been completed. ❌ 41 New Failures, 3 Unrelated Failures, 39 Unclassified FailuresAs of commit 707e35f with merge base ed65b12 ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
| `executorch.kernels.quantized` loads a plugin that carries its own copy of the same | ||
| kernels, and the runtime stops when the same operator is registered twice: |
There was a problem hiding this comment.
can we make this as clean as other .so files i.e. registered only once, and not part of the cpython extension lib?
| library torch loads at export time, because each side registers into a table the | ||
| other never reads. Loading both into one process does abort on the second | ||
| registration, so what this check enforces for that component is one owner among the | ||
| runtime libraries, not the absence of the export copy. Excusing every component | ||
| would disarm the check where duplication is a real fault: two of these libraries | ||
| defined the backend registry symbols in one released wheel and not in the release | ||
| before it, so the duplication this catches does happen. |
| with torch.no_grad(): | ||
| expected = model(*example) | ||
|
|
||
| if mode == "quantized": |
digantdesai
left a comment
There was a problem hiding this comment.
ok with duplicate symbols but we should see if we can do something about it.
Good question, and I dug into it rather than guessing. The short answer is that the duplicate is Measured on the installed wheel, there are two separate libraries that both register the same They are not two copies of the same thing by accident. They serve different sides:
So the kernels are already out of the cpython extension, which is the part your comment asks about. Making that "registered only once" means having the export-time plugin reuse the runtime component What I have done here is make the conflict impossible to hit by accident: the component is held out |
A quantized model uses smaller numbers than a normal one, so the tensors take less memory.
Running one needs the quantized operator kernels.
The only copy the wheel shipped is the one torch loads to export a model, which a C++ application
cannot use. Such an application links the runtime, loads a quantized model, and the model fails at
run time with a missing operator, which looks like a model problem rather than a packaging one.
Build the quantized kernels as their own shared library and name it as a CMake component, the same
way the other kernel sets are named.
The wheel now ships
lib/libexecutorch_kernels_quantized.so.Note that the wheel also ships a second copy of these kernels, inside the library torch loads when
you export a model. That copy is built into the plugin rather than resolved from the shared library,
so a process holding both registers the same operators twice, and the runtime treats that as fatal:
This affects only a process that does both, for example an application that embeds a Python
interpreter. A plain C++ application can link the component freely.
Because of that, this is the one component
EXECUTORCH_LIBRARIESdoes not include, so anapplication that links whatever the package offers cannot end up in that position without asking.
A consumer that wants the quantized kernels names the component, or on CMake older than 3.28, where
no component targets exist, links
EXECUTORCH_QUANTIZED_KERNELS_LIBRARYas well. That variable isnow populated on both CMake routes, so a consumer that adopts the older-CMake recipe and later
upgrades keeps the library on their link line instead of silently losing it.
Built the wheel, installed it into a clean environment, and:
quantization step (measured worst difference 0.0048 against a tolerance of 0.02).
executorch::kernels_quantized, ran the same program, and gotthe same output as Python, byte for byte.
the shipped library and the export plugin aborts in either load order.
collides: the CPU kernels, the delegate, the thread pool, the profiler and the runtime all
coexist with both the extension and the export plugin.
EXECUTORCH_LIBRARIESdoes not depend on the quantized library whilestill depending on the CPU kernels, on CMake 3.28 and on real CMake 3.24. A new check asserts
this, and it fails on the previous behaviour.
EXECUTORCH_QUANTIZED_KERNELS_LIBRARYresolves to the shipped library on both the modern-CMakeroute (as the imported target) and the pre-3.28 route (as a file path).
the wheel enables these kernels unconditionally, so their absence is a regression rather than a
configuration to tolerate, and both the ownership table and the C++ check previously treated it as
an acceptable state and reported coverage they had not run.
Ran on Linux x86_64 and aarch64.