Add and validate support for a Hugging Face model architecture in the Optimum Intel OpenVINO backend, including exporter configuration, patching, repository tests, and documentation.
日本語の概要は準備中です。原文の説明を表示しています。
Create a small random Hugging Face model that preserves the original architecture and can serve as a local Optimum Intel repository test fixture.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Create a tiny random model for the requested architecture without loading the original model weights. The artifact must preserve the real architecture and execute the requested task path.
If repairing a tiny model after a validation or benchmark failure, start from the exact failing local artifact, creation script, traceback, and task reproducer. Repair that artifact and rerun the same command; do not generate a different model and use its success as evidence for the failed one.
Download or load configuration and lightweight code/processor assets only. Record:
model_type, architectures, auto_map, and Transformers metadata;Determine and record the Transformers version or version range supported by the original model, especially for trust-remote-code architectures. Create and validate the tiny model using a compatible version; do not silently substitute another model implementation because the active Transformers version lacks the architecture.
Inspect precision fields under every relevant configuration level. Remote
models may use dtype, torch_dtype, or both, including separate values in
vision, text, audio, or projector sub-configs.
Create create_tiny_model.py in the designated working directory. It must:
Define maximum parameter-count and model-memory budgets, then verify both after construction so reducing layer count cannot be offset by widening other dimensions. Estimate weight memory from each parameter's element count and element size, and reject a candidate that exceeds either budget.
Do not reuse a cache merely because config.json and a weight file exist.
Before returning it, validate a cache-format/version marker and all critical
configuration invariants, including architecture identity, dimensions,
special tokens, processor assets, and nested precision fields. Rebuild the
cache when the generator logic or required invariants change.
Keep construction logic easy to adapt into
_create_tiny_<model_type>_model() in tests/openvino/utils_tests.py.
Compare the original and tiny configurations. Preserve:
model_type and architectures;When deliberately forcing a test model to float32, update and verify every
effective precision field used by the remote configuration (dtype and/or
torch_dtype, including nested sub-configs). Reload the saved model and check
its actual parameter dtypes; editing an ignored config key is not sufficient.
Never rename the model type, substitute a nearby architecture, or remove a component merely to make export pass.
If the tiny model's model_type, architectures, or task-relevant component
identity differs from the original model, stop and report the mismatch. A tiny
model that executes successfully through another architecture is not a valid
fixture.
Reload the saved directory through its documented Transformers or pipeline API and execute the requested task.
Follow the task-specific tiny-model validation and output-validity instructions
supplied for <task>.
If task execution fails, repair the violated configuration invariant, recreate the model, and rerun it. Loading, saving, or a forward pass alone is not success.
The final validation evidence must load the exact output directory returned by the creator, execute the requested task, and include the command and output. Do not validate one directory and return a different cached or previously generated artifact.
Report the output directory, script path, configuration comparison, parameter count, exact task execution command, output, and any dependency or remote-code constraints.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Add and validate support for a Hugging Face model architecture in the Optimum Intel OpenVINO backend, including exporter configuration, patching, repository tests, and documentation.
日本語の概要は準備中です。原文の説明を表示しています。