Skip to content

Pull requests: gittensor-ai-lab/sparkinfer

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

perf(qwen36): hoist the Q6_K unpack out of the DFlash head multi-row loop (+3.9% DFLASH_TPS) 32k-context UI-only: strongest measured context in sparkinfer eval area:kernels subsystem (emission weight 0.42) dflash-merge-first round winner: biggest verified DFlash speedup — auto-merge candidate eval:infra-error sparkinfer auto-eval verdict: infra-error eval-dflash:XS sparkinfer DFlash vs-main speed tier: XS re-evaluate Winner merged — rebase onto main; bot re-evaluates on push test-on-5090 Maintainer-approved to evaluate on RTX 5090 (greenlight)
#666 opened Jul 29, 2026 by nickmopen Contributor Loading…
1 task done
server: emit GPU ttft/generation/decode_tps in stream usage area:runtime subsystem (emission weight 0.26) hold Maintainer override: never auto-merge this PR
#570 opened Jul 21, 2026 by ai-hpc Member Loading…
1 of 4 tasks
ProTip! Updated in the last three days: updated:>2026-07-30.