-
Notifications
You must be signed in to change notification settings - Fork 219
Pull requests: waybarrios/vllm-mlx
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(mllm): fail closed on unknown drafter batch support
#755
opened Aug 31, 2026 by
Thump604
Collaborator
Loading…
fix(mllm): load registered MTP drafters safely
#752
opened Aug 31, 2026 by
Thump604
Collaborator
Loading…
fix(mllm): reconcile stateful MTP batches atomically
#754
opened Aug 31, 2026 by
Thump604
Collaborator
Loading…
fix(api): reject unavailable Outlines backend
#753
opened Aug 31, 2026 by
Thump604
Collaborator
Loading…
chore: remove dead _extract_media_from_messages helper
#751
opened Aug 31, 2026 by
djacobsmeyer
Contributor
Loading…
fix(mllm): populate vllm_mlx_engine_steps_executed for MLLM-routed models
#749
opened Aug 30, 2026 by
djacobsmeyer
Contributor
Loading…
5 tasks done
fix(shutdown): exit the serve process before MLX's thread-local teardown segfaults
#745
opened Aug 30, 2026 by
sbayer2
Loading…
fix(qwen4-exp): keep text inference on mlx-vlm path
#741
opened Aug 28, 2026 by
Thump604
Collaborator
Loading…
Fix streaming token loss at --stream-interval>1, DSA CacheList SSD spill, bf16 upcast on mlx 0.31.x, and forced-<think> reasoning
#740
opened Aug 27, 2026 by
anisoptera
Loading…
fix(registry): populate /metrics engine-state gauges in registry mode
#739
opened Aug 27, 2026 by
djacobsmeyer
Contributor
Loading…
fix(scheduler): guard MTP install against mlx-lm >= 0.31 BatchGenerator
#737
opened Aug 27, 2026 by
azamamirza
Loading…
test(simple-engine): key stream ownership by Thread, not ident
#734
opened Aug 26, 2026 by
janhilgard
Collaborator
Loading…
fix(video): video requests silently dropped, then broken on mlx-vlm 0.6
#733
opened Aug 24, 2026 by
sbayer2
Loading…
feat(api): surface prefix-cache reuse via usage.prompt_tokens_details.cached_tokens
#732
opened Aug 24, 2026 by
djacobsmeyer
Contributor
Loading…
fix(mllm): make prefix-cache rewind gate aware of non-KV hybrid leaves
#731
opened Aug 24, 2026 by
djacobsmeyer
Contributor
Loading…
fix(cli): honor --prefill-step-size with continuous batching
#729
opened Aug 23, 2026 by
Thump604
Collaborator
Loading…
fix(mllm): keep media requests out of prefix cache
#728
opened Aug 23, 2026 by
Thump604
Collaborator
Loading…
fix(reasoning): GLM-4.7 injects <think> in the prompt — stream thinking as reasoning_content, not content
#725
opened Aug 23, 2026 by
Rune-Humborstad
Loading…
fix(mllm): dispatch assistant-drafter loading by architecture instead of hardcoding Gemma 4
#724
opened Aug 22, 2026 by
brandy975
Contributor
Loading…
Add capacity envelope benchmark artifacts
#720
opened Aug 21, 2026 by
Thump604
Collaborator
Loading…
Previous Next
ProTip!
Updated in the last three days: updated:>2026-08-28.