# Receipt Recognition System 

We aim to build a receipt recognition and information extraction system that could answer questions based on information on a given receipt photo. We want it to not only be able to correctly extract out the exact numeric value on the receipt when such task is given, but also to answer more complex questions that involve some math calculation.

## Loading and testing out the vision model

In [4]:
# !pip install accelerate>=0.26.0

In [8]:
from PIL import Image
import torch
from transformers import AutoModelForVision2Seq, AutoProcessor, BitsAndBytesConfig

In [6]:
access_token = "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"
# Hugging Face model id
model_id = "meta-llama/Llama-3.2-11B-Vision-Instruct"

# Load model and tokenizer
model = AutoModelForVision2Seq.from_pretrained(
    model_id,
    device_map="auto",
    torch_dtype=torch.bfloat16,
    token=access_token
)

processor = AutoProcessor.from_pretrained(
    model_id,
    token=access_token
)

2024-11-25 03:25:26.786551: E tensorflow/compiler/xla/stream_executor/cuda/cuda_dnn.cc:9342] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
2024-11-25 03:25:26.786587: E tensorflow/compiler/xla/stream_executor/cuda/cuda_fft.cc:609] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
2024-11-25 03:25:26.786617: E tensorflow/compiler/xla/stream_executor/cuda/cuda_blas.cc:1518] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
2024-11-25 03:25:26.794093: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
The model weights are not tied. Please use the `ti

Loading checkpoint shards:   0%|          | 0/5 [00:00<?, ?it/s]

In [12]:
image_path = "receipt-recognition/receipt_00000.png"
image = Image.open(image_path)

messages = [
    {"role": "user", "content": [
        {"type": "image"},
        {"type": "text", "text": "Read the following receipt and extract the same type of information (Store Name, Date, Items, Tax, and Total)."}
    ]}
]
input_text = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(
    image,
    input_text,
    add_special_tokens=False,
    return_tensors="pt"
).to(model.device)

output = model.generate(**inputs, max_new_tokens=200)
print(processor.decode(output[0]))

<|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>Read the following receipt and extract the same type of information (Store Name, Date, Items, Tax, and Total).<|eot_id|><|start_header_id|>assistant<|end_header_id|>

**Receipt Information**

* **Store Name:** Not visible on the receipt
* **Date:** Not visible on the receipt
* **Items:**
	+ Nasi Campur Bali
	+ Bok Bengil Nasi
	+ MilkShake Starwb
	+ Ice Lemon Tea
	+ Nasi Ayam Dewata
	+ Free Ice Tea
	+ Organic Green Sa
	+ Ice Tea
	+ Ice Orange
	+ Ayam Suir Bali
	+ Tahu Goreng
	+ Tempe Goreng
	+ Tahu Telor Asin
	+ Nasi Goreng Samb
	+ Bok Panggang Sam
	+ Ayam Sambal Hija
	+ Hot Tea
	+ Ice Kopi
	+ Tahu Telor Asin
	+ Free Ice Tea
	+ Bebek Street
	+ Ice Tea Tawar
* **Tax:** Not visible on the receipt



## Evaluation

We have created a jsonl file with 11 evaluation benchmark questions. From the following result, we could see that the model have a high accuracy on pure information extraction question, which means that the model can recognize the text on receipts very well. But when it answer questions that involve math calculation, such as unit price and total payment after a tax is applied, it has a very bad accuracy using the current basic prompting. We will try out more sophisticated prompting skills to help the model make a better decision on those types of more complex problems.

In [5]:
import json

file_path = 'qa.jsonl'

with open(file_path, 'r') as file:
    for line in file:
        data = json.loads(line)
        
        print(f"Image path: {data['image_path']}")
        print(f"question: {data['question']}")
        print(f"Expected Answer: {data['expected']}")
        print() 

Image path: receipt-recognition/receipt_00000.png
question: What is the total payment amount?
Expected Answer: 1591600

Image path: receipt-recognition/receipt_00000.png
question: How much was spent on hot tea in total ?
Expected Answer: 44000

Image path: receipt-recognition/receipt_00000.png
question: What is the unit price of hot tea?
Expected Answer: 22000

Image path: receipt-recognition/receipt_00000.png
question: How much does the ice tea cost?
Expected Answer: 18000

Image path: receipt-recognition/receipt_00000.png
question: How much does the free ice tea cost?
Expected Answer: 0

Image path: receipt-recognition/receipt_00000.png
question: How many Bbk Panggang Sam was bought?
Expected Answer: 3

Image path: receipt-recognition/receipt_00000.png
question: How much is the service fee?
Expected Answer: 10950

Image path: receipt-recognition/receipt_00000.png
question: If a 10% tax is applied on the subtotal, what will the total payment now be?
Expected Answer: 1726200

Image pat

In [12]:
import json

file_path = 'qa.jsonl'

def format_question_prompt(question):
    return f"Read the following receipt. {question}. Respond in a sigle numeric value or 'not found' if it doesn't exist in the given image."

with open(file_path, 'r') as file:
    for line in file:
        data = json.loads(line)
        
        print(f"Image path: {data['image_path']}")
        print(f"question: {data['question']}")
        print(f"Expected Answer: {data['expected']}")
        
        image_path = data['image_path']
        image = Image.open(image_path)

        messages = [
            {"role": "user", "content": [
                {"type": "image"},
                {"type": "text", "text": format_question_prompt(data['question'])}
            ]}
        ]
        input_text = processor.apply_chat_template(messages, add_generation_prompt=True)
        inputs = processor(
            image,
            input_text,
            add_special_tokens=False,
            return_tensors="pt"
        ).to(model.device)

        output = model.generate(**inputs, max_new_tokens=200)
        print(f"Model Output: {processor.decode(output[0])}")
        print() 

Image path: receipt_00000.png
question: What is the total payment amount?
Expected Answer: 1591600
Model Output: <|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>Read the following receipt. What is the total payment amount?. Respond in a sigle numeric value or 'not found' if it doesn't exist in the given image.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

1,591,600.<|eot_id|>

Image path: receipt_00000.png
question: How much was spent on hot tea in total ?
Expected Answer: 44000
Model Output: <|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>Read the following receipt. How much was spent on hot tea in total?. Respond in a sigle numeric value or 'not found' if it doesn't exist in the given image.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

44,000.<|eot_id|>

Image path: receipt_00000.png
question: What is the unit price of hot tea?
Expected Answer: 22000
Model Output: <|begin_of_text|><|start_header_id|>user<|end_header_id|>

<