vllm.config.engram ¶
Classes:
-
EngramConfig–Configuration for Engram embedding storage and sharding.
Functions:
-
model_has_engram_layers–Whether the model carries n-gram embedding layers.
EngramConfig ¶
Configuration for Engram embedding storage and sharding.
Methods:
-
compute_hash–Hash settings that affect embedding execution and graph structure.
-
get_parallel_size–Derive the embedding group size from the parallel configuration.
-
verify_load_config–Shared tables require a loader that invokes parameter weight callbacks.
-
verify_model_config–Reject Engram configuration for models without n-gram embeddings.
-
verify_parallel_config–Reject unsupported embedding parallel topologies.
Attributes:
-
cpu_offload(bool) –Store embedding weights in pinned CPU memory for UVA lookup.
-
dp_shared_memory(bool) –Share CPU-offloaded embedding weights between co-located
-
embedding_across_dp(bool) –Shard embeddings across TP and all DP ranks when enabled.
Source code in vllm/config/engram.py
cpu_offload = Field(default_factory=_default_cpu_offload) class-attribute instance-attribute ¶
Store embedding weights in pinned CPU memory for UVA lookup. Defaults to VLLM_PLE_CPU_OFFLOAD, which is enabled by default. An explicit value takes precedence over the environment variable.
dp_shared_memory = False class-attribute instance-attribute ¶
Share CPU-offloaded embedding weights between co-located DP replicas. Each node stores one copy of every TP shard, reducing host memory without per-step Engram DP collectives. Requires sufficient /dev/shm capacity and a shared IPC namespace.
embedding_across_dp = False class-attribute instance-attribute ¶
Shard embeddings across TP and all DP ranks when enabled. Otherwise, each DP rank has a separate TP-sharded embedding replica.
compute_hash() ¶
get_parallel_size(parallel_config) ¶
Derive the embedding group size from the parallel configuration.
Source code in vllm/config/engram.py
verify_load_config(load_config) ¶
Shared tables require a loader that invokes parameter weight callbacks.
Source code in vllm/config/engram.py
verify_model_config(model_config) ¶
Reject Engram configuration for models without n-gram embeddings.
Source code in vllm/config/engram.py
verify_parallel_config(parallel_config) ¶
Reject unsupported embedding parallel topologies.
Source code in vllm/config/engram.py
model_has_engram_layers(model_config) ¶
Whether the model carries n-gram embedding layers.