AlpinDale a4cbcfe59f feat: disable logprob serialization to CPU for spec decode преди 5 месеца
..
__init__.py 9d81716bfd [v0.5.3] Release Candidate (#388) преди 8 месеца
batch_expansion.py 2c653a2268 fix: make speculative decoding work with per-request seed преди 5 месеца
draft_model_runner.py a4cbcfe59f feat: disable logprob serialization to CPU for spec decode преди 5 месеца
interfaces.py 3a53ff1e01 fix: raise an error for no draft token case when draft_tp>1 преди 5 месеца
medusa_worker.py 16dff9babc chore: enable bonus token in spec decoding for KV cache based models преди 5 месеца
metrics.py 2ebb37d1ee update time since last collection for AsyncMetricsCollector преди 5 месеца
mlp_speculator_worker.py 16dff9babc chore: enable bonus token in spec decoding for KV cache based models преди 5 месеца
multi_step_worker.py dd18c5042c move prepare_inputs to the GPU (#596) преди 5 месеца
ngram_worker.py 16dff9babc chore: enable bonus token in spec decoding for KV cache based models преди 5 месеца
proposer_worker_base.py d638dc592d fix: some minor typing issues in spec decode преди 5 месеца
smaller_tp_proposer_worker.py 16dff9babc chore: enable bonus token in spec decoding for KV cache based models преди 5 месеца
spec_decode_worker.py a4cbcfe59f feat: disable logprob serialization to CPU for spec decode преди 5 месеца
target_model_runner.py a4cbcfe59f feat: disable logprob serialization to CPU for spec decode преди 5 месеца
top1_proposer.py 3a53ff1e01 fix: raise an error for no draft token case when draft_tp>1 преди 5 месеца
util.py a4cbcfe59f feat: disable logprob serialization to CPU for spec decode преди 5 месеца