Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Partitioned LongBench

Most performance benchmarks for inference frameworks test using small context sizes. For example Llama 2 on Amazon SageMaker a Benchmark, has a mean of 233 and max of 1024 tokens (measured using the LLaMA 2 SentencePiece Tokenizer).

The performance characteristics of inference framework deployments using long contexts of 2,000 - 4,--- or mroe tokens are not commonly measured.

The LongBench dataset contains some long examples with query use cases and is available on GitHub at https://github.com/THUDM/LongBench.

This repository partitions LongBench samples by context length, which may be applied to prompt and request templates to form benchmark performance and load test datasets. An example dataset for LLaMA 2 is provided.

Reference LLaMA 2 Prompt Template

This is based on LLaMA 2 generation.py for a QASum task, with a well-known extension adding the instruction after the final [/INST], see LLaMA Recipes - Prompt Llama 2.

<s>[INST] <<SYS>>
{{ system_prompt }}
<</SYS>>

```
{{ context }}
```

Question: {{ question }}

[/INST]
Answer:

where system_prompt is adapted from LangChain's prompt hub.

You are an assistant for question-answering tasks. Use the following pieces of \
retrieved context in the section demarcated by "```" to answer the question. \
If you don't know the answer just say that you do not know. Use three sentences
maximum and keep the answer concise.

So the final prompt is:

<s>[INST] <<SYS>>
You are an assistant for question-answering tasks. Use the following pieces of \
retrieved context in the section demarcated by "```" to answer the question. \
If you don't know the answer just say that you don't know. Use three sentences
maximum and keep the answer concise.
<</SYS>>

```
{{ context }}
```

Question: {{ question }}

[/INST]
Answer:

Using the LLaMA 2 Tokenizer, the above prompt, when context and question are empty, is 98 tokens.

Large Files

The dataset included here is over 100 MB so you will need to use git-lfs.

About

https://github.com/THUDM/LongBench data partitioned into Retrieval Augmented Generation (RAG) segments for performance testing.

Resources

Stars

0 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages