Repository navigation
Vidify 2.0 Release Notes
Overview
Vidify 2.0 brings powerful upgrades to our AI-powered Chrome extension that lets users search for objects and words within YouTube videos. With the addition of real-time object detection and a dynamic "table of contents" feature, this release goes beyond transcript-only search—marking a major milestone toward smart, efficient video navigation.
What’s New in 2.0
Visual Object Detection (YOLOv8 + Grounding DINO)
- Detects key objects directly from video frames
- Users can view an auto-generated table of contents based on visual objects
- Uses YOLOv8 for speed and Grounding DINO for zero-shot, niche object support
- Achieves ~70% accuracy on COCO dataset
Object-Based Table of Contents
- Scene-by-scene breakdown of detected objects
- Clickable timestamps for jumping directly to key visual moments
YOLO Model Integration
- Integrated into backend pipeline for scalable object detection
- Ready for GPU deployment for higher performance
- Supports both standard and custom detection terms
Other Improvements
- Refined UI for toggling between object and transcript search
- Backend API supports dual processing routes (transcript + object detection)
- Faster API response and improved data parsing
- More precise indexing of both captions and object metadata
Installation & Requirements
- Chrome v88+
- Python 3.9+
- Internet connection
- Optional: GPU for local model training or inference
Known Limitations
- Object detection currently only available for YouTube videos with accessible content
- No auto-seek preview yet for object detection timestamps
- May miss some domain-specific or rare objects due to generalization
Future Roadmap
- Custom training for improved niche detection
- Compatibility with Netflix, Amazon Prime, and other streaming services
- Video summarization based on objects and transcript context
- Full-seek integration for object detection matches
Thanks for using Vidify!