-
Notifications
You must be signed in to change notification settings - Fork 2.5k
Pull requests: JustVugg/colibri
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
The VRAM expert tier could only be fed by disk reads
#877
opened Aug 7, 2026 by
ZacharyZcR
Contributor
Loading…
Avoid reserving DSpark cache for unsupported DeepSeek-V4 checkpoints
#875
opened Aug 7, 2026 by
waizuichougou
Loading…
A CUDA fallback is a property of the call, not of the tensor
bug
Difetto verificato nel codice
cuda
Backend CUDA/NVIDIA
#849
opened Aug 6, 2026 by
ZacharyZcR
Contributor
Loading…
docs: M1 Ultra fmt=2 Metal perf report + benchmarks row + fmt4→fmt2 converter
benchmark
Datapoint di misurazione hardware
docs
Documentazione
metal
Backend Metal/Apple
#836
opened Aug 5, 2026 by
boblaublaw
Loading…
inkling: ring-buffer KV cache for sliding-window layers (~10x less KV memory at long context)
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#830
opened Aug 4, 2026 by
dpanelli
Loading…
fix(metal): int4-g64 (fmt=4) models never used the GPU — enable MoE experts + fused attention
bug
Difetto verificato nel codice
metal
Backend Metal/Apple
needs-rebase
Confligge, serve rebase dell'autore
#829
opened Aug 4, 2026 by
aaristov
Loading…
engine: notice when a tensor's format silently disables the fused Metal decode path
enhancement
New feature or request
metal
Backend Metal/Apple
#827
opened Aug 4, 2026 by
monotophic
Contributor
Loading…
Feat/dstorage transport
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
#825
opened Aug 4, 2026 by
khalilswdp
Contributor
Loading…
4 of 5 tasks
managed runner, deterministic benchmark, and causal evidence (PR3 draft of #377 split)
discussion
Proposta / discussione aperta, non un task
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
ramdisk: headless planning, staging, mounts, recovery + tokenized CLI (PR2 of #377 split)
discussion
Proposta / discussione aperta, non un task
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
engine: RAMMAP, NUMA, and telemetry (PR1 of #377 split)
discussion
Proposta / discussione aperta, non un task
enhancement
New feature or request
performance
Velocità / tok-s / ottimizzazioni
Add a registry for the 214 environment variables, and check the environment against it
enhancement
New feature or request
#800
opened Aug 3, 2026 by
ZacharyZcR
Contributor
Loading…
Dead code, de-duplication, repo-wide clang-format, and lint gates in CI
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#798
opened Aug 3, 2026 by
ZacharyZcR
Contributor
Loading…
Metal (Apple GPU) backend for Kimi K3
feature
Nuova funzionalità
metal
Backend Metal/Apple
model-support
Supporto a nuovi modelli
#790
opened Aug 2, 2026 by
RDouglasSharp
Contributor
Loading…
[WIP] Vulkan stats
enhancement
New feature or request
vulkan
Backend Vulkan/AMD
#789
opened Aug 2, 2026 by
Neppord
Loading…
5 tasks
feature: Hy3 support
discussion
Proposta / discussione aperta, non un task
model-support
Supporto a nuovi modelli
#775
opened Aug 2, 2026 by
ErikTromp
Loading…
5 tasks done
DeepSeek V4: the CUDA kernels, as a tier rather than an engine
discussion
Proposta / discussione aperta, non un task
model-support
Supporto a nuovi modelli
#772
opened Aug 2, 2026 by
ZacharyZcR
Contributor
Loading…
Vulkan MoE GEMV backend for integrated/AMD GPUs (draft, complements #418)
enhancement
New feature or request
feature
Nuova funzionalità
vulkan
Backend Vulkan/AMD
feat(qwen36): CUDA VRAM expert tier — heat-based placement across GPUs via the shared CUDA backend
cuda
Backend CUDA/NVIDIA
model-support
Supporto a nuovi modelli
feat(qwen36): Qwen3.6-35B-A3B engine (CPU): hybrid Gated Attention + Gated DeltaNet + streaming MoE
model-support
Supporto a nuovi modelli
#712
opened Jul 30, 2026 by
kreuzzelg
Contributor
Loading…
feat(win): fix silent CPU fallback, launcher suite, DirectStorage expert loads
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#670
opened Jul 28, 2026 by
khalilswdp
Contributor
Loading…
8 tasks done
feat(core): model-architecture seam, chat templates, text antiprompt
enhancement
New feature or request
needs-rebase
Confligge, serve rebase dell'autore
#667
opened Jul 28, 2026 by
khalilswdp
Contributor
Loading…
5 tasks done
QLoRA training path: fine-tune GLM-5.2 (744B) in 64 GB RAM
enhancement
New feature or request
feature
Nuova funzionalità
#626
opened Jul 26, 2026 by
pavolbauer
Loading…
5 tasks done
WIP: MiniMax-M3 support — GQA + MSA block-sparse attention, o200k tokenizer, converter (follow-up to #418)
model-support
Supporto a nuovi modelli
#601
opened Jul 24, 2026 by
steve-m
Contributor
Loading…
5 tasks
Metal fmt=4 grouped-int4 decode: attention + routed experts (#585)
metal
Backend Metal/Apple
needs-rebase
Confligge, serve rebase dell'autore
#587
opened Jul 24, 2026 by
RDouglasSharp
Contributor
Loading…
Previous Next
ProTip!
Adding no:label will show everything without a label.