v1.0.1
Pre-release
Pre-release
Version 1.0.1 (pre-release) is an almost bug-free version for v1 Overhaul.
New Features:
- StateCacheManager and CPUAttention supports for fp16 precision on ARM v82
- CPUAttention fused tiled operator for paged attention (still suffer from kind of MAUVE loss under current quantization)
- Add 9 samplers: "greedy", "temperature", "topK", "topP", "minP", "tfs", "typical", "penalty", "penalize_ngram"
- Add perplexity test unit on wikitext2
Latest Release will be ready if the master branch is up to date to that of alibaba/MNN.
- Implement BeamSearchSampler (including StateCacheManager supports)
- Research on KV cache quantization algorithm, and implement at least 1.
- Resolve conflicting implementation of KVCache, sampler, and single model in alibaba/MNN to prepare for merge, and then merge
- Move the control follow (all_seq_len_ and gen_seq_len_) to sampler. Llm is changed into a simple inference module.
- Support composite sampler, e.g., allow penalty+topK+topP, etc.
- Move the prompt library to an independent class.