vllm.v1.worker.gpu.sample.watermark ¶
Functions:
-
philox_gumbel_block_argmax–Apply the detector-compatible PRF to one vocabulary block.
-
repeated_context_mask–Return, per row, whether the row's context already occurred in its history.
_repeated_context_mask_cpu(all_token_ids, req_indices, prompt_lens, total_lens, contexts, max_history=None, include_prompt=False, skip_partial_context=False) ¶
Reference implementation; the parity tests check the Triton kernel against it.
Source code in vllm/v1/worker/gpu/sample/watermark.py
philox_gumbel_block_argmax(logits, mask, block_idx, contexts_row_ptr, key_0, key_1, CONTEXT_WIDTH, BLOCK_SIZE) ¶
Apply the detector-compatible PRF to one vocabulary block.
Source code in vllm/v1/worker/gpu/sample/watermark.py
repeated_context_mask(all_token_ids, req_indices, prompt_lens, total_lens, contexts, max_history=None, include_prompt=False, skip_partial_context=False) ¶
Return, per row, whether the row's context already occurred in its history.
Parameters:
-
(all_token_ids¶Tensor) –[max_num_reqs, max_model_len]token ids of every request. -
(req_indices¶Tensor) –Request slot per sampled row; -1 marks a padding row, which is reported as not repeated.
-
(prompt_lens¶Tensor) –Prompt length per request slot.
-
(total_lens¶Tensor) –Prompt plus generated length per request slot.
-
(contexts¶Tensor) –[num_rows, context_width]context per row, padded with -1 before the start of the scanned history. -
(max_history¶int | None, default:None) –Number of most recent history positions searched, or
Nonefor all of them. The window compared at each position reachescontext_widthtokens further back. -
(include_prompt¶bool, default:False) –Search the prompt as well as the generated tokens.
-
(skip_partial_context¶bool, default:False) –Mark contexts containing start padding so they use ordinary sampling.