First, we need to install the Hugging Face Transformers library, which gives us access to state-of-the-art speech and language models. If you haven’t installed it yet, run:

```
pip install transformers torch
```

Now, let’s import the necessary libraries and set up our device for computation. If you have a GPU, the code will automatically use it for better performance!

## Transcribing mp3 file to numpy requires the FFmpeg. Follow the steps in this link to have it installed. [(How to install FFmpeg in windows)](https://www.wikihow.com/Install-FFmpeg-on-Windows)

In [1]:
from transformers import pipeline,AutoModelForCausalLM, AutoTokenizer
import torch
from IPython.display import  clear_output


device = 'cuda' if torch.cuda.is_available() else 'cpu'

# SPEECH-TO-TEXT WITH WHISPER

# If MP4 (or other video format) convert to MP3

In [2]:
import subprocess

def convert_mp4_to_mp3(input_file, output_file):
    """Convert an MP4 file to MP3 using ffmpeg."""
    try:
        command = [
            "ffmpeg",
            "-i", input_file,    # Input file
            "-vn",               # No video
            "-acodec", "libmp3lame",  # MP3 codec
            "-b:a", "192k",      # Audio bitrate
            output_file          # Output file
        ]
        subprocess.run(command, check=True)
        print(f"Conversion successful: {output_file}")
    except subprocess.CalledProcessError as e:
        print(f"Error during conversion: {e}")

# Example usage
convert_mp4_to_mp3("2025-03-21 10-15-55.mp4", "meeting.mp3")


ffmpeg version 4.4.2-0ubuntu0.22.04.1 Copyright (c) 2000-2021 the FFmpeg developers
  built with gcc 11 (Ubuntu 11.2.0-19ubuntu1)
  configuration: --prefix=/usr --extra-version=0ubuntu0.22.04.1 --toolchain=hardened --libdir=/usr/lib/x86_64-linux-gnu --incdir=/usr/include/x86_64-linux-gnu --arch=amd64 --enable-gpl --disable-stripping --enable-gnutls --enable-ladspa --enable-libaom --enable-libass --enable-libbluray --enable-libbs2b --enable-libcaca --enable-libcdio --enable-libcodec2 --enable-libdav1d --enable-libflite --enable-libfontconfig --enable-libfreetype --enable-libfribidi --enable-libgme --enable-libgsm --enable-libjack --enable-libmp3lame --enable-libmysofa --enable-libopenjpeg --enable-libopenmpt --enable-libopus --enable-libpulse --enable-librabbitmq --enable-librubberband --enable-libshine --enable-libsnappy --enable-libsoxr --enable-libspeex --enable-libsrt --enable-libssh --enable-libtheora --enable-libtwolame --enable-libvidstab --enable-libvorbis --enable-libvpx --enab

Conversion successful: meeting.mp3


size=   15908kB time=00:11:18.67 bitrate= 192.0kbits/s speed= 141x    
video:0kB audio:15908kB subtitle:0kB other streams:0kB global headers:0kB muxing overhead: 0.004445%


Next, we’ll use OpenAI’s Whisper model to transcribe an audio file. We’re using the whisper-small model for efficiency, but you can switch to whisper-large for better accuracy.

In [5]:
pipe  = pipeline("automatic-speech-recognition",
                    "openai/whisper-small", 
                    chunk_length_s=30,
                    stride_length_s=5,
                    return_timestamps=True,
                    device=device, 
                    generate_kwargs = {"language": 'english', "task": "transcribe"}) 

Device set to use cuda


Now, let’s process an audio file. The model will transcribe and translate it into English."

In [6]:
transcription = pipe("meeting.mp3" )
#Once the transcription is complete, we’ll format the text with timestamps for better readability.
formatted_lyrics = ""
for line in transcription['chunks']:
    text = line["text"]
    ts = line["timestamp"]
    formatted_lyrics += f"{ts}-->{text}\n"

print(formatted_lyrics.strip())

You have passed task=transcribe, but also have set `forced_decoder_ids` to [[1, None], [2, 50359]] which creates a conflict. `forced_decoder_ids` will be ignored in favor of task=transcribe.
The attention mask is not set and cannot be inferred from input because pad token is same as eos token. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.


(0.0, 107.0)--> Thank you. What if AI could help you find life's biggest question straight from Bhagavad Gita? In this video, I'll show you how to build an AI system that retrieves verses from the Gita and answers any live questions you ask. For example, here is a sample question, why does God allow suffering and evil in the world? This is a generated answer by AI in Bhagavad Gita, Lord Krishna explains, allows suffering and evil, but it's the natural order of the universe, sound and so forth. You can also get it in the... In the Bhagavad Gita, Lord Krishna explains that God does not allow suffering and evil.
(107.0, 159.0)--> You can get that in the format of speaking as well. So, in this video, you will learn how to combine retrieval of Maintain Generation or Ragn shot with the timeless Ujjdam of Gita. So you can create an AI powered spiritual guide for you.
(246.24, 408.64)--> So can AI really provide meaningful answers? Stay till the end to see it in action. Here is a sample questi

In [7]:
# Let’s also save the transcription to a text file so we can use it later!"
with open("transcription.txt", "w", encoding="utf-8") as file:
    file.write(formatted_lyrics.strip())

print("Transcription saved to transcription.txt")

Transcription saved to transcription.txt


# SUMMARIZING WITH LLAMA
Now, let’s take it a step further! We’ll use Meta’s LLaMA model to summarize the transcript. For this, we need the Llama-3.2-3B-Instruct model.

In [8]:
DEFAULT_MODEL = "meta-llama/Llama-3.2-3B-Instruct"


model = AutoModelForCausalLM.from_pretrained(
    DEFAULT_MODEL,
    torch_dtype=torch.bfloat16,
    use_safetensors=True,
    device_map=device,
)

tokenizer = AutoTokenizer.from_pretrained(DEFAULT_MODEL, use_safetensors=True)
tokenizer.pad_token_id = tokenizer.eos_token_id

Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

We’ll define a conversation prompt instructing LLaMA to summarize the meeting transcript in simple English.

In [9]:
conversation = [
    {"role": "system", "content": '''You are a master in YouTube Script Writing and Information Delivering without making a viewer feel bored. I am working on a YouTube Script for a Video [Title]. I need a complete skeleton structure for it with all the points included, don’t miss any. In the skeleton structure, each point should include What should be Included in this point, What does the Viewer expect from this point (not in terms of feelings, in terms of information included and presentation) and How should this information be presented in a flow. Don’t forget to include examples of each point that give me an idea on how to write the script myself. I’m writing this script in a human conversational tone so keep that in mind while writing your examples. If there is any need of providing any reference, study results, mechanism, science backed techniques, facts or anything for any point in any part of the script to make it more informative, mention that in that particular point not at the end. Now using all your expertise write me a skeleton structure with every point included and some examples for each of them.'''},
    {"role": "user", "content": f'''{formatted_lyrics.strip()}'''},
]

prompt = tokenizer.apply_chat_template(conversation, tokenize=False)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
# print(prompt)

with torch.no_grad():
    output = model.generate(
        **inputs,
        do_sample=True,
        max_new_tokens=1000
    )

processed_text = tokenizer.decode(output[0][len(inputs.input_ids[0]):], skip_special_tokens=True)

print(processed_text)

Setting `pad_token_id` to `eos_token_id`:128001 for open-end generation.


assistant

Based on the provided script, I've created a skeleton structure for your YouTube video on "How to build an AI system that retrieves verses from the Bhagavad Gita and answers live questions." Here's the skeleton structure with each point, what the viewer expects from each point, and how to present the information in a flow:

**Introduction (0.0, 107.0)**

* What to include: Introduce the concept of the video, explain the goal of building an AI system that retrieves verses from the Bhagavad Gita and answers live questions.
* What the viewer expects: An overview of the topic, understanding of the purpose of the video.
* How to present: Start with a hook to grab the viewer's attention, provide a brief introduction to the topic, and explain the goal of the video.

Example:
"Welcome to our video on how to build an AI system that retrieves verses from the Bhagavad Gita and answers live questions. In this video, we'll show you how to combine retrieval of text with the timeless wisdo

The provided text appears to be a transcript of a meeting between multiple individuals, including Maximilian Moro, Carlos, and others. The meeting seems to be discussing various topics, including:

1. **Control and inspection of components**: The group discusses the importance of controlling and inspecting components, particularly in the context of a map or a system. They talk about the need for a camera to capture images of the components and the importance of aligning them.
2. **Camera setup and testing**: The group discusses the setup and testing of cameras, including the use of different types of cameras and the importance of testing for parallelism and alignment.
3. **Image processing and analysis**: The group talks about the need for image processing and analysis, including the use of software to detect defects and anomalies in the images.
4. **Feasibility studies**: The group discusses the importance of feasibility studies, including the creation of a set of images for each control and the testing of these images to determine the conditions of success.
5. **Camera placement and ergonomics**: The group talks about the importance of camera placement and ergonomics, including the need to position cameras to capture the best images and to reduce the risk of errors.
6. **Camera calibration and maintenance**: The group discusses the importance of camera calibration and maintenance, including the need to change batteries and perform regular checks on the cameras.
7. **Image quality and resolution**: The group talks about the importance of image quality and resolution, including the need to capture high-quality images to ensure accurate inspection and analysis.
8. **Labeling and annotation**: The group discusses the importance of labeling and annotating images, including the need to clearly label and identify components and defects.

Some specific points mentioned in the transcript include:

* The need to use a camera with a good resolution to capture high-quality images
* The importance of aligning images to ensure accurate inspection and analysis
* The need to test cameras for parallelism and alignment
* The importance of feasibility studies to determine the conditions of success for each control
* The need to position cameras to capture the best images and reduce the risk of errors
* The importance of regular camera calibration and maintenance
* The need to clearly label and annotate images to ensure accurate inspection and analysis

Overall, the meeting appears to be focused on improving the inspection and analysis process using cameras and image processing techniques, and ensuring that the cameras are properly set up and maintained to capture high-quality images.