vllm.models.glm5next.common.model ¶
_dequant_fp8_block(weight_fp8, scale_inv, block_size=128) ¶
Dequantize a block-FP8 (e4m3) weight with per-block scale to BF16.
Unlike scaled_dequantize this tolerates a non-divisible (partial last block) shape by zero-padding to a multiple of block_size before the scale broadcast and trimming back afterwards (e.g. kv_a_proj_with_mqa is 576 rows = 4*128 + 64).
Source code in vllm/models/glm5next/common/model.py
_try_load_fp8_attn_proj(name, tensor, buf, params_dict, loaded_params, kv_a_pad_size) ¶
Dequantize FP8 q_a_proj / kv_a_proj_with_mqa / o_proj to BF16 on load.
The FP8 checkpoint stores these as block-FP8 (weight + weight_scale_inv), but the model holds them in BF16 (fused_qkv_a_proj is always BF16 via DeepSeekV2FusedQkvAProjLinear; o_proj is excluded by modules_to_not_convert). When the model target is BF16 (no weight_scale_inv param) we dequantize; otherwise we return False so the normal stacked/direct path loads the FP8 tensor as-is.
Source code in vllm/models/glm5next/common/model.py
1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 | |