Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

18 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation


β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•—   β–ˆβ–ˆβ•—β–ˆβ–ˆβ•—  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•—β–ˆβ–ˆβ–ˆβ•—   β–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ•—
β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘  β•šβ•β•β–ˆβ–ˆβ•”β•β•β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ•— β–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•”β–ˆβ–ˆβ–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β–ˆβ–ˆβ–ˆβ–ˆβ•”β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘
β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘     β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•”β•β•β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘
     β–ˆβ–ˆβ•‘ β•šβ•β• β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—β–ˆβ–ˆβ•‘   β–ˆβ–ˆβ•‘β–ˆβ–ˆβ•‘ β•šβ•β• β–ˆβ–ˆβ•‘β•šβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•”β•β–ˆβ–ˆβ•‘  β–ˆβ–ˆβ•‘β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ•—
     β•šβ•β•     β•šβ•β• β•šβ•β•β•β•β•β• β•šβ•β•β•β•β•β•β•β•šβ•β•   β•šβ•β•β•šβ•β•     β•šβ•β• β•šβ•β•β•β•β•β• β•šβ•β•β•β•β•β• β•šβ•β•  β•šβ•β•β•šβ•β•β•β•β•β•β•
                           AI  PIPELINE
        Video Transcription with Audio Gap Reconstruction

An enterprise-grade, serverless system that reconstructs missing audio gaps using multimodal visual evidence β€” cursor tracking, UI element detection, OCR, and GPT-4o reasoning.


Python Azure Functions Azure Video Indexer GPT-4o OpenCV


Build Cloud Internship



πŸ“‹ Table of Contents

  1. Overview
  2. The Business Problem
  3. Key Differentiators & Features
  4. Architecture Deep Dive
  5. Core AI Models & APIs Used
  6. Repository Structure
  7. System Requirements & Prerequisites
  8. Local Development Setup
  9. Configuration & Environment Variables
  10. Running the System Locally
  11. Output Format & Results
  12. Cloud Deployment to Azure
  13. Testing & Validation
  14. Troubleshooting & Known Issues
  15. Acknowledgements

πŸ”­ Overview

The Multimodal AI Pipeline is a state-of-the-art event reconstruction system built natively on Azure Durable Functions. It is designed to solve the problem of missing or corrupted audio in screen recordings (e.g., training videos, software demos, meeting recordings).

Instead of failing silently or producing blank transcripts during audio dropouts, this system intelligently pivots to visual evidence. It analyzes frame-by-frame cursor movements, clicks, Optical Character Recognition (OCR) data, and workflow changes, merging them with any available transcribed speech. Finally, it uses Azure OpenAI's GPT-4o to generate a cohesive, human-readable narrative documenting exactly what happened in the video.

The pipeline is fully serverless, highly parallelized via a fan-out/fan-in design, and scales infinitely using Azure Blob Storage triggers.


🧩 The Business Problem

Standard transcription services (like Whisper, AWS Transcribe, or traditional STT engines) make one fatal assumption: The audio is the only signal of value.

When an employee records a critical 10-minute software walkthrough, and their microphone cuts out for 2 minutes in the middle, traditional tools simply return [BLANK AUDIO]. The context of those 2 minutes is permanently lost.

This pipeline treats screen recordings as a multi-channel event stream. If the audio channel fails, the visual channel is deeply analyzed to answer the question: "What was the user attempting to accomplish?"

What the System Discovers Visually:

Visual Signal Extracted Meaning
πŸ–±οΈ Cursor Velocity Drops Intent to interact with an element
πŸ–±οΈ Optical Flow Anomalies Mouse clicks and drag actions
πŸ“œ Azure Vision OCR Global context, statuses, ID numbers, rich form data
πŸ”„ Structural Scene Changes Screen transitions, modal opens, and workflow steps

✨ Key Differentiators & Features

1. Azure Durable Functions Orchestration (No Complex Infrastructure)
There are no Celery queues, Redis caches, or complex Docker Swarms to maintain. Azure Durable Functions natively handles distributed state, orchestrator lifetimes, and retry policies.

2. Fan-Out/Fan-In Extreme Parallelism
The orchestrator launches the heavy ActivityAudioPipeline and ActivityVisualPipeline completely in parallel. A 10-minute video processes almost twice as fast because audio transcription and frame extraction happen simultaneously in isolated subprocesses.

3. Fully Dynamic Prompts & Self-Structuring Data
We don't hardcode rules to look for "User IDs" or "Names." The Azure Computer Vision (Read API) harvests all key-value pairs on screen. GPT-4o is dynamically prompted to ingest whatever rich data it receives and intuitively weave it into the narration.

4. Isolated Ephemeral Workspaces
Azure Functions run multiple tasks concurrently in the same Python worker process. To prevent catastrophic race conditions with "Current Working Directories," each pipeline activity spawns a heavily isolated subprocess.run() environment mapped to temporary /tmp folders.


πŸ›οΈ Architecture Deep Dive

High-Level Design (HLD)

The pipeline uses an event-driven, trigger-based architecture. A simple video upload initiates a massively parallel workflow.

graph TD
    A[Blob Storage: input-videos] -->|BlobTrigger Event| B(video_upload_starter)
    B -->|Spawns Instance| C{PipelineOrchestrator}
    
    C -->|Fan-Out Async| D[ActivityAudioPipeline]
    C -->|Fan-Out Async| E[ActivityVisualPipeline]
    
    D -->|Polls AVI REST API| F(audio_events.json)
    E -->|OpenCV + Vision OCR API| G(visual_events.json & extracted_details.json)
    
    F -->|Fan-In Wait| H[ActivityFusion]
    G -->|Fan-In Wait| H
    
    H -->|Constructs Master Timeline| I[Temporal Alignment Engine]
    I -->|Prompts GPT-4o| J[Document Generator]
    J -->|Saves Markdown| K[Blob Storage: output-docs]
Loading

Low-Level Component Workflows

1. Audio Processing Pipeline (audio_main.py)

  • Isolation: Runs inside an ephemeral /tmp directory via Python subprocess.
  • Action: Submits the video to Azure Video Indexer.
  • Resilience: Implements a polling loop with try/except exponential backoff to survive transient network resets ([WinError 10054]) over long 10-minute transcription windows.
  • Output: Transcripts, speaker diarization, and exact word-level timestamps saved to state JSON.

2. Visual Processing Pipeline (main.py)

  • Keyframe Extraction: Uses OpenCV to detect scene changes (Structural Similarity/Frame Differencing) to avoid processing identical frames.
  • Cursor Tracking: Analyzes optical flow. A sudden deceleration of the cursor typically indicates a "click" or intent.
  • OCR Analysis: Sends keyframes to Azure AI Vision Read API. Uses intelligent heuristics (bounding box geometry) to pair labels with values.
  • Output: Structured UI events and rich key-value data extracted from the screen.

3. Temporal Fusion & Narration (fusion_main.py & document_generator.py)

  • Alignment: Aligns the millisecond timestamps of Audio events and Visual events onto a single unified master clock.
  • Context Blocks: Chunks the timeline into semantic "scenes."
  • LLM Narration: Passes the scenes and unstructured visual data to GPT-4o with strict system prompts enforcing professional, third-person documentation phrasing without hallucinating unsupported facts.

🧠 Core AI Models & APIs Used

Subsystem Cloud Provider Specific Service / Model Role in Pipeline
Audio STT Microsoft Azure Azure Video Indexer Transcription & Speaker Diarization
Vision OCR Microsoft Azure Azure AI Vision (Read API) Extracting text, forms, and UI states
Reasoning Microsoft Azure Azure OpenAI (GPT-4o) Fusing multi-modal data into narration
Computer Vision Open Source OpenCV (cv2) Optical Flow, Frame Differencing

πŸ“ Repository Structure

multimodal-pipeline/
β”‚
β”œβ”€β”€ πŸ“„ README.md                    ← You are here! Detailed documentation.
β”œβ”€β”€ πŸ“„ local.settings.json          ← Azure Functions config (⚠️ NEVER COMMIT THIS)
β”œβ”€β”€ πŸ“„ host.json                    ← Azure Functions runtime configuration
β”œβ”€β”€ πŸ“„ requirements.txt             ← Python dependencies
β”œβ”€β”€ πŸ“„ function_app.py              ← ⚑ Serverless Entry Point (Orchestrator & Triggers)
β”‚
β”œβ”€β”€ πŸ“ audio_stream/                ← πŸ”Š Audio Subsystem
β”‚   └── πŸ“„ avi_client.py            ← Azure Video Indexer REST API wrapper (w/ Retries)
β”‚
β”œβ”€β”€ πŸ“ visual_stream/               ← πŸ‘οΈ Visual Subsystem
β”‚   β”œβ”€β”€ πŸ“„ frame_extractor.py       ← OpenCV Keyframe Extraction logic
β”‚   β”œβ”€β”€ πŸ“„ cursor_tracker.py        ← Optical Flow & Click detection
β”‚   β”œβ”€β”€ πŸ“„ ocr_engine.py            ← Azure Vision Read API Client
β”‚   └── πŸ“„ ocr_detail_extractor.py  ← Key-Value Pair bounding box correlator
β”‚
β”œβ”€β”€ πŸ“ fusion/                      ← 🧠 Brain Subsystem
β”‚   └── πŸ“„ document_generator.py    ← GPT-4o dynamic prompt assembly
β”‚
β”œβ”€β”€ πŸ“„ main.py                      ← CLI wrapper for Visual Pipeline Subprocess
β”œβ”€β”€ πŸ“„ audio_main.py                ← CLI wrapper for Audio Pipeline Subprocess
β”œβ”€β”€ πŸ“„ fusion_main.py               ← CLI wrapper for Fusion Pipeline Subprocess
β”œβ”€β”€ πŸ“„ azure_clients.py             ← Shared Azure SDK initializers (OpenAI, Storage)
└── πŸ“„ utils.py                     ← Shared helpers (JSON I/O, file paths)

βš™οΈ System Requirements & Prerequisites

To run and develop this pipeline locally, ensure your environment meets the following baseline:

Local Machine Prerequisites

  1. OS: Windows 10/11, macOS (Apple Silicon/Intel), or Ubuntu 20.04+.
  2. Python: Version 3.10, 3.11, or 3.12 (Azure Functions V4 runtime does not yet fully support 3.13 in all regions).
  3. Azure Functions Core Tools: Version 4.x (Installation Guide)
  4. Git: For version control.

Cloud Azure Prerequisites

You must have an Active Azure Subscription with the following deployed resources:

  1. Azure Storage Account (v2)
  2. Azure Video Indexer Account
  3. Azure AI Vision (Computer Vision resource, standard tier)
  4. Azure OpenAI Service (With a gpt-4o model deployment)

πŸ’» Local Development Setup

Follow these exact steps to get the pipeline running on your local machine.

1. Clone & Bootstrap

# Clone the repository
git clone https://github.com/Yash12930/testcase.git
cd testcase

# Create a virtual environment
python -m venv venv

# Activate the virtual environment
# On Windows:
.\venv\Scripts\activate
# On macOS/Linux:
source venv/bin/activate

# Install required dependencies
pip install -r requirements.txt

2. Configure Local Azure Storage (Optional but Recommended)

If you do not want to test directly against cloud blob storage, you can use Azurite (The Azure Storage Emulator).

npm install -g azurite
azurite --silent --location c:\azurite --debug c:\azurite\debug.log

Note: The system requires actual Azure API keys for Video Indexer and OpenAI regardless of where blobs are stored.


πŸ” Configuration & Environment Variables

Azure Functions relies heavily on the local.settings.json file for local development. This file is git-ignored for security.

Create a local.settings.json file in the root of your project:

{
  "IsEncrypted": false,
  "Values": {
    "FUNCTIONS_WORKER_RUNTIME": "python",
    "AzureWebJobsFeatureFlags": "EnableWorkerIndexing",
    
    "AzureWebJobsStorage": "DefaultEndpointsProtocol=https;AccountName=YOUR_STORAGE_ACCOUNT;AccountKey=YOUR_KEY;EndpointSuffix=core.windows.net",
    
    "AZURE_STORAGE_CONNECTION_STRING": "DefaultEndpointsProtocol=https;AccountName=YOUR_STORAGE_ACCOUNT;AccountKey=YOUR_KEY;EndpointSuffix=core.windows.net",
    
    "AVI_ACCOUNT_ID": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx",
    "AVI_API_KEY": "your_video_indexer_subscription_key",
    "AVI_LOCATION": "trial",
    
    "AZURE_VISION_ENDPOINT": "https://your-vision-resource.cognitiveservices.azure.com/",
    "AZURE_VISION_KEY": "your_vision_subscription_key",
    
    "AZURE_OPENAI_ENDPOINT": "https://your-openai-resource.openai.azure.com/",
    "AZURE_OPENAI_API_KEY": "your_openai_subscription_key",
    "AZURE_OPENAI_DEPLOYMENT_NAME": "gpt-4o"
  }
}

Variable Breakdown:

  • AzureWebJobsStorage: Used by Azure Durable Functions to maintain state and queue tasks.
  • AZURE_STORAGE_CONNECTION_STRING: Used by our custom scripts to download/upload videos to the input-videos and output-docs containers.

▢️ Running the System Locally

With your local.settings.json configured, starting the pipeline is as simple as launching the Azure Functions host.

# Ensure your virtual environment is activated
func start

How to Trigger a Run:

  1. Open Azure Storage Explorer or the Azure Portal.
  2. Navigate to your Storage Account and ensure the containers input-videos, temp-state, and output-docs exist.
  3. Upload any .mp4 file into the input-videos container.
  4. Watch the local terminal! You will see:
    • The BlobTrigger firing immediately.
    • The Orchestrator starting.
    • [Activity-Visual] and [Activity-Audio] spinning up concurrent subprocesses.
    • Polling logs from Azure Video Indexer.
    • Final Fusion logs indicating document generation.

πŸ“ Output Format & Results

Once the pipeline completes, a Markdown (.md) file will automatically appear in the output-docs container.

Sample Output Extract:

# Process Workflow Document
**Generated By:** Multimodal AI Pipeline
**Video Source:** `demo_recording_04.mp4`

## Timeline

* **00:00 - 00:45 [AUDIO]**
  *Speaker 1:* "Alright, today I'm going to show you how to provision a new user in the active directory system."

* **00:45 - 01:20 [VISUAL INFERENCE]**
  *The user navigated to the 'Admin Panel'. They clicked the 'Create New User' button. Based on screen data, the 'Employee ID' field was populated with '#99281A' and the 'Department' dropdown was set to 'Finance'. The user then confirmed the creation.*

* **01:20 - 01:40 [AUDIO]**
  *Speaker 1:* "As you can see, it takes a few seconds to sync. Once it's done, we can exit."

Notice how the Visual Inference block intelligently extracted 'Employee ID' and 'Department' dynamically without being explicitly hardcoded to look for them, filling the exact gap where the user stopped speaking!


☁️ Cloud Deployment to Azure

Deploying this pipeline to an Azure Function App (Consumption or Premium Plan) is straightforward.

1. Create the Function App Resource

Ensure you create a Python Function App on a Linux OS plan.

2. Deploy via Azure CLI

# Log in to Azure
az login

# Publish the code to your Function App
func azure functionapp publish <YourFunctionAppName>

3. Migrate Environment Variables

The settings inside local.settings.json are not deployed automatically. You must manually add them to the Function App's "Environment Variables" / "Application Settings" via the Azure Portal or CLI:

az functionapp config appsettings set --name <YourFunctionAppName> --resource-group <YourResourceGroup> --settings AZURE_OPENAI_API_KEY="your-key" AVI_API_KEY="your-key" ...

πŸ§ͺ Testing & Validation

Validating Architecture Concurrency

To ensure the Azure CWD subprocess isolation is working correctly:

  1. Upload a 6-minute+ video.
  2. Monitor the func start logs.
  3. Ensure both [Activity-Audio] and [Activity-Visual] show overlapping logs without throwing FileNotFoundError or cv2.imwrite silent failures on outputs/frames.

Subprocess Testing

You can manually test individual pipeline scripts without running the whole Durable Functions orchestrator:

python main.py path/to/local/video.mp4
python audio_main.py path/to/local/video.mp4

🚨 Troubleshooting & Known Issues

Issue Root Cause Solution
ConnectionResetError [WinError 10054] during Audio Pipeline Azure Video Indexer dropped the idle polling connection on videos > 5 mins. Fixed natively. avi_client.py uses try/except backoff to seamlessly reconnect.
Invalid frame path: outputs\frames\... in OCR Global CWD race condition between parallel Azure Functions activities. Fixed natively. function_app.py uses subprocess.run(cwd=tmpdir) to strictly isolate memory execution.
BlobTrigger fails to fire AzureWebJobsStorage connection string is invalid, or the storage account lacks permissions. Verify your connection string in local.settings.json. Ensure you aren't using an expired SAS token.
OpenCV NoneType errors on video load Video file path is incorrect, or the blob did not download completely to /tmp. Verify storage container names. Ensure tmpdir disk space is not exhausted.
Missing Structured Details GPT-4o output is too generic or hallucinating. Ensure Azure Vision Read API is returning valid bounding boxes in extracted_details.json.

πŸ™ Acknowledgements

  • Microsoft Azure Serverless Team β€” For Azure Durable Functions, enabling highly scalable stateful logic.
  • OpenAI & Azure Cognitive Services β€” For GPT-4o and Vision APIs, which make complex visual reasoning possible.
  • OpenCV Community β€” For reliable, battle-tested computer vision implementations.
  • Tech Mahindra β€” For providing the internship opportunity and challenging architectural problems to solve.

Multimodal AI Pipeline Β· Tech Mahindra Internship 2026
Confidential β€” Internal Use Only

Built with too much caffeine and genuine curiosity about what happens at the intersection of vision, language, and audio.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages