Skip to content

v1.8.0 - Advanced AI Features

Choose a tag to compare

@doublegate doublegate released this 02 Oct 03:18
· 78 commits to main since this release

🤖 Advanced AI Features

This release transforms WebScrape-TUI into an AI-powered content intelligence platform with automatic tagging, named entity recognition, keyword extraction, and semantic similarity matching capabilities.

✨ Major Features

AI-Powered Auto-Tagging

  • Intelligent Tag Generation: Automatically generate tags from article content using AI analysis
  • TF-IDF Extraction: Advanced keyword extraction with confidence scoring
  • Tag Deduplication: Automatic cleaning and deduplication of generated tags
  • Configurable Parameters: Adjust max tags and confidence thresholds
  • AITaggingManager Class: 125 lines of robust tag generation logic
  • Keyboard Shortcut: Ctrl+Shift+T

Use Cases:

  • Automatic content categorization without manual tagging
  • Discover hidden topics and themes in your content
  • Build tag taxonomies from large article collections

Named Entity Recognition (NER)

  • spaCy Integration: Industry-standard NLP with en_core_web_sm model
  • Entity Types: Extract people (PERSON), organizations (ORG), locations (GPE/LOC), dates (DATE), products (PRODUCT)
  • Categorized Display: Organized modal showing entities by type
  • Entity Deduplication: Automatic removal of duplicate entities
  • EntityRecognitionManager Class: 82 lines of entity extraction and categorization
  • Keyboard Shortcut: Ctrl+Shift+E

Use Cases:

  • Identify key people, organizations, and locations in articles
  • Research and fact-checking workflows
  • Building knowledge graphs from scraped content
  • Quick content overview without reading full articles

Content Similarity Matching

  • Semantic Embeddings: SentenceTransformer (all-MiniLM-L6-v2) for deep content understanding
  • Similarity Detection: Find related articles using cosine similarity
  • Configurable Threshold: Adjust similarity cutoff (default: 0.5)
  • Top-K Selection: Retrieve N most similar articles
  • ContentSimilarityManager Class: 108 lines of similarity computation
  • Keyboard Shortcut: Ctrl+Shift+R

Use Cases:

  • Discover related content and connections
  • Duplicate detection and deduplication
  • Content clustering and organization
  • Research topic exploration

Keyword Extraction

  • TF-IDF Scoring: Statistical keyword importance ranking
  • Title Boosting: Prioritize keywords appearing in titles
  • NLTK Stopword Filtering: Remove common words for better results
  • Frequency Analysis: Multi-word phrase support
  • KeywordExtractionManager Class: 99 lines of keyword extraction logic
  • Keyboard Shortcut: Ctrl+Shift+K

Use Cases:

  • Quick content overview and summarization
  • SEO and content analysis
  • Topic trending and analysis
  • Building content indexes

Multi-Level Summarization

  • Three Summary Levels: Brief (1-2 sentences), Detailed (paragraph), Comprehensive (full summary)
  • AI Provider Integration: Works with Gemini, OpenAI, and Claude
  • Extractive & Abstractive: TF-IDF sentence selection + AI generation
  • Progressive Detail: Adapt summary length to your needs
  • MultiLevelSummarizationManager Class: 78 lines of summarization logic

Use Cases:

  • Generate executive summaries for stakeholders
  • Create different summary versions for different audiences
  • Quick content triage and filtering
  • Documentation and report generation

🔧 Technical Implementation

New AI Manager Classes (492 lines total)

  1. AITaggingManager (scrapetui.py:2537-2661)

    • generate_tags_from_content(): TF-IDF-based tag extraction
    • suggest_tags_with_ai(): AI-powered tag suggestions
    • Tokenization with NLTK
    • Confidence scoring and filtering
  2. EntityRecognitionManager (scrapetui.py:2663-2744)

    • load_spacy_model(): Lazy loading of spaCy model
    • extract_entities(): Entity extraction with categorization
    • Support for 5 entity types
    • Deduplication and ordering
  3. ContentSimilarityManager (scrapetui.py:2745-2852)

    • load_model(): Lazy loading of SentenceTransformer
    • calculate_similarity(): Cosine similarity computation
    • find_similar_articles(): Top-K similar article retrieval
    • Efficient embedding generation
  4. KeywordExtractionManager (scrapetui.py:2853-2951)

    • extract_keywords(): TF-IDF keyword extraction
    • extract_key_phrases(): Multi-word phrase extraction
    • Title keyword boosting
    • NLTK stopword integration
  5. MultiLevelSummarizationManager (scrapetui.py:2952-3042)

    • generate_one_sentence_summary(): Ultra-concise summaries
    • generate_extractive_summary(): TF-IDF sentence selection
    • AI provider abstraction
    • Progressive detail levels

UI Integration (240 lines)

New Keybindings (scrapetui.py:4755-4758):

  • Ctrl+Shift+T: Auto-tag article
  • Ctrl+Shift+E: Extract entities
  • Ctrl+Shift+K: Extract keywords
  • Ctrl+Shift+R: Find similar articles

Action Methods (scrapetui.py:5602-5644):

  • action_auto_tag(): Triggers auto-tagging workflow
  • action_extract_entities(): Triggers entity extraction
  • action_extract_keywords(): Triggers keyword extraction
  • action_find_similar(): Triggers similarity search

Async Workers (scrapetui.py:5976-6207):

  • _auto_tag_worker(): Background tag generation (46 lines)
  • _extract_entities_worker(): Background entity extraction (79 lines)
  • _extract_keywords_worker(): Background keyword extraction (44 lines)
  • _find_similar_worker(): Background similarity search (58 lines)

📦 Dependencies

6 New NLP/ML Libraries:

spacy>=3.7.0                    # Natural language processing
sentence-transformers>=2.2.0    # Semantic embeddings
nltk>=3.8.0                     # Text processing toolkit
scikit-learn>=1.3.0             # TF-IDF vectorization
scipy>=1.11.0                   # Cosine distance calculation

Additional Setup Required:

# Install dependencies
pip install -r requirements.txt

# Download spaCy language model (required for NER)
python -m spacy download en_core_web_sm

# NLTK data (downloads automatically on first run)

🧪 Testing

28 New Tests across 5 comprehensive test classes (515 lines):

  • TestAITagging (6 tests):

    • Basic tag generation from content
    • Empty/short content handling
    • Tag limit enforcement
    • Tag deduplication
    • AI provider failure handling
    • Tag cleaning and normalization
  • TestEntityRecognition (6 tests):

    • Basic entity extraction
    • Empty content handling
    • No entities scenario
    • Entity deduplication
    • Multiple entity types
    • spaCy model error handling
  • TestContentSimilarity (6 tests):

    • Basic similarity calculation
    • Threshold filtering
    • Top-K selection
    • Empty database handling
    • Single article edge case
    • Model loading error handling
  • TestKeywordExtraction (6 tests):

    • Basic keyword extraction
    • Empty content handling
    • Top-N selection
    • Frequency-based scoring
    • Title keyword boosting
    • NLTK error handling
  • TestMultiLevelSummarization (4 tests):

    • Level-specific summary generation
    • All levels generation
    • Empty content handling
    • AI provider error handling

Total Test Suite: 194 tests (100% pass rate)

📖 Documentation

Comprehensive Updates:

  1. README.md:

    • AI Integration section enhanced with v1.8.0 features
    • 6 new dependencies documented
    • 4 new keyboard shortcuts added to shortcuts table
    • Test count updated to 194 tests
    • Recent updates section highlighting v1.8.0
  2. CHANGELOG.md:

    • Complete v1.8.0 entry (175 lines)
    • Detailed feature descriptions with use cases
    • Technical implementation details
    • Testing information and results
    • Files modified with line numbers
  3. docs/ROADMAP.md:

    • Current status updated to v1.8.0
    • v1.8.0 features moved to completed section
    • v1.9.0 planned features outlined
  4. docs/PROJECT-STATUS.md:

    • Version updated to v1.8.0
    • Quick stats updated (194 tests, 17 dependencies)
    • Current development phase documented
    • v1.8.0 feature completeness table added

🎯 Use Cases

Content Research

  • Entity Extraction: Quickly identify key people, organizations, and locations in articles
  • Keyword Analysis: Extract main topics and themes without reading full content
  • Similarity Matching: Discover related articles and build topic clusters

Content Management

  • Auto-Tagging: Automatically categorize and organize large article collections
  • Duplicate Detection: Find and merge similar/duplicate articles
  • Smart Organization: Build tag taxonomies and content hierarchies

Analysis & Reporting

  • Multi-Level Summaries: Generate summaries for different audiences (executives, analysts, researchers)
  • Entity Mapping: Build knowledge graphs of people, organizations, and locations
  • Trend Analysis: Extract keywords to identify trending topics over time

Research Workflows

  • Topic Exploration: Use similarity matching to explore related content
  • Fact Checking: Extract entities and verify information
  • Literature Review: Auto-tag and organize research articles

📈 Quick Start

# Install new dependencies
pip install spacy sentence-transformers nltk scikit-learn scipy

# Download spaCy model (required for entity recognition)
python -m spacy download en_core_web_sm

# Or reinstall all dependencies
pip install -r requirements.txt
python -m spacy download en_core_web_sm

# Run application
python scrapetui.py

# Try AI features (select an article first):
Press Ctrl+Shift+T  # Auto-tag article
Press Ctrl+Shift+E  # Extract entities
Press Ctrl+Shift+K  # Extract keywords
Press Ctrl+Shift+R  # Find similar articles

🔗 Documentation Links

🚀 What's Next?

v1.9.0 (Q1 2026) - Smart Categorization & Topic Modeling

  • Topic modeling with LDA and NMF
  • Advanced entity relationship mapping
  • Knowledge graph construction
  • Advanced similarity clustering
  • Enhanced content analysis

See the Roadmap for full details.

📊 Project Stats

  • Lines of Code: ~6,500 in main file (+1,000 from v1.7.0)
  • Test Coverage: 194 tests across 13 test suites (100% pass rate)
  • Features: 65+ major capabilities
  • Documentation: 120KB+ comprehensive guides
  • Dependencies: 17 production, 2 development

🎉 Highlights

This release brings cutting-edge AI capabilities to WebScrape-TUI:

  • ✅ Automatic content understanding with AI
  • ✅ Named entity recognition for information extraction
  • ✅ Semantic similarity for content discovery
  • ✅ Intelligent keyword extraction
  • ✅ Multi-level summarization
  • ✅ 28 new comprehensive tests
  • ✅ Production-ready NLP/ML integration

Full Changelog: v1.7.0...v1.8.0

🤖 Generated with Claude Code