-
Notifications
You must be signed in to change notification settings - Fork 17
Contributors
This fork is developed and tested by people who contribute code, reviews, benchmarks, documentation, and access to hardware.
Main developer and maintainer. Leads the fork's architecture, implementation, integration, releases, documentation, and performance validation.
Major code and testing contributor. Merged contributions include:
- Qwen3.8 MTP and DSpark validation and fixes (PR #3)
- Native quantized KV FlashAttention (PR #4)
- CUDA build and configuration fixes (PR #34)
- Benchmark controls for KV placement and workspace testing (PR #70)
- Correct multi-stream scheduler range copies (PR #77)
- Broader architecture coverage for rollback, quantization, and parallel sequences (PR #94)
PiggiDragon also maintains ongoing work on host-KV transport, multi-GPU placement, speculative decoding, and native quantized FlashAttention. See all submitted pull requests.
The following people donated machines or remote machine access to GenerelSchwerz for testing that the project could not cover locally.
| Contributor | Hardware | What it enabled |
|---|---|---|
| Add contributor | CPU, GPU, RAM, and operating system | Validation or benchmark work |
Only add people here with their permission. Describe the hardware and the work it enabled so the credit remains useful.
Code, testing, reproducible benchmark reports, documentation corrections, and platform-specific validation are all useful. Open a pull request against the relevant fork branch and include enough evidence for someone else to reproduce the result.
GenerelSchwerz llama.cpp
- Home
- Discord community
- Contributors
- Complete feature index
- Hardware setup guides
- Owner-verified evidence
- Notable runs
- Benchmark comparison showcase
- BeeLlama Main
- Llama Main
- Llama Dev
- MoE Cache
Feature groups
- Memory placement and workspace
- Validation and diagnostics
- CUDA MoE cache and helpers
- Grouped MoE drafting
Source branches