Releases: Embedded-AI-Systems/MNN-LLM
Releases · Embedded-AI-Systems/MNN-LLM
Release list
v1.0.1
Version 1.0.1 (pre-release) is an almost bug-free version for v1 Overhaul.
New Features:
- StateCacheManager and CPUAttention supports for fp16 precision on ARM v82
- CPUAttention fused tiled operator for paged attention (still suffer from kind of MAUVE loss under current quantization)
- Add 9 samplers: "greedy", "temperature", "topK", "topP", "minP", "tfs", "typical", "penalty", "penalize_ngram"
- Add perplexity test unit on wikitext2
Latest Release will be ready if the master branch is up to date to that of alibaba/MNN.
- Implement BeamSearchSampler (including StateCacheManager supports)
- Research on KV cache quantization algorithm, and implement at least 1.
- Resolve conflicting implementation of KVCache, sampler, and single model in alibaba/MNN to prepare for merge, and then merge
- Move the control follow (all_seq_len_ and gen_seq_len_) to sampler. Llm is changed into a simple inference module.
- Support composite sampler, e.g., allow penalty+topK+topP, etc.
- Move the prompt library to an independent class.
v1.0.0
MNN-LLM v1.0.0
New Features:
- Individual Sampler (greedy, temperature, top K, top P, min P )
- StateCacheManager (Paged Attention, Fused CPU Attention op)
qwen1.5-4B-chat model in MNN format with fused attention operator is appended in the Release binary.