Skip to content

Repository files navigation

Jabbah

Parse, query, and clean real-world HTML in pure Ruby.

Gem version CI status Ruby 3.2 or newer MIT license

Features · Installation · Quick start · Sanitization · Article extraction


Jabbah is a dependency-free HTML parser for static documents and reader applications. It tolerates broken markup, builds a traversable tree, supports a useful CSS selector subset, and provides conservative sanitization. It never fetches pages or executes JavaScript. The name comes from ν Scorpii and the Arabic jabha, “forehead.”

Features

  • Document and fragment parsing with implicit element closing
  • Traversal, cloning, serialization, and common CSS selectors
  • UTF-8, BOM, and declared-charset decoding
  • Sanitization profiles for documents, feeds, and mail
  • Remote-image blocking and readability-style article extraction
  • No runtime dependencies or native extensions

Installation

Add gem "jabbah" to your Gemfile and run bundle install, or install directly:

gem install jabbah

Requires Ruby 3.2 or newer.

Quick start

require "jabbah"

document = Jabbah.parse('<main><h1>Hello</h1><p class="lead">Welcome.</p></main>')
puts document.at("main h1").text          # => Hello
puts document.search("p.lead").first.text # => Welcome.
puts document.to_html

Jabbah.parse returns a Jabbah::Document. Use Jabbah.fragment for fragments without document wrappers. #at returns the first match and #search returns every match.

Sanitization

safe = Jabbah::Sanitize.clean(
  document,
  profile: :feed,
  base_url: "https://example.com/"
)
puts safe.to_html
puts "Blocked images: #{safe.blocked_count}"

The :docs, :feed, and :mail profiles remove active elements and unsafe URLs. Remote images are blocked by default, with their original URL retained in data-blocked-src; cid: images are preserved. Sanitization returns a clone, leaving the original document unchanged.

Article extraction

article = Jabbah::Extract.article(safe)
puts article[:title] if article
puts article[:content].text if article

Extraction is heuristic and may return nil when no article-like content is found. See the tree and serialization decision for the underlying design.

Development

bundle install
bundle exec rake test

License

MIT

About

Dependency-free pure Ruby HTML parser with selectors, sanitization, and article extraction.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages