-
Notifications
You must be signed in to change notification settings - Fork 0
Adding a Site
Set up the repository first, as CONTRIBUTING describes.
Look for its file under lib/recipe_scrapers/sites/, or ask in code:
RecipeScrapers::Registry.for("https://www.example.com/recipes/pancakes")
# => nil when the site is not supportedChoose a recipe with several steps, because a recipe that reads as a single step usually means the steps were not read correctly. If the site groups its ingredients on some recipes, pick one with groups.
Most recipe sites publish schema.org markup, and then the site needs no code. Read the page without the site check:
require "net/http"
require "recipe_scrapers"
url = "https://www.example.com/recipes/pancakes"
html = Net::HTTP.get(URI(url))
pp RecipeScrapers.parse(html, url: url, supported_only: false).to_hCompare the output with the page in a browser: the title, every ingredient line, every step and the ingredient groups.
The file lives at lib/recipe_scrapers/sites/<top-level domain>/<first label>.rb, and
RecipeScrapers::SitePath.for("example.com") returns that path, here com/example.
If the output was right, the file is one line:
# frozen_string_literal: true
RecipeScrapers.register "example.com"If some fields were wrong or missing, declare where they are. See Declarations.
# frozen_string_literal: true
RecipeScrapers.register "example.com" do
ingredients rows: ".recipe-ingredients li"
instructions rows: ".recipe-steps li"
endIf the markup is there but has to be read in its own way, write a Scraper subclass. See
Declarations.
A site that serves the same recipes on several domains registers them together:
RecipeScrapers.register "bonviveur.com", also: ["bonviveur.es"]The spec lives at the same path as the site file, here spec/sites/com/example_spec.rb:
# frozen_string_literal: true
RSpec.describe "example.com" do
subject(:recipe) { scrape_cassette("com/example", url: "https://www.example.com/recipes/pancakes") }
it "reads the title" do
expect(recipe.title).to eq("Pancakes")
end
it "reads every ingredient line" do
expect(recipe.ingredients).to eq(["200 g flour", "2 eggs", "300 ml milk"])
end
it "reads every instruction step" do
expect(recipe.instructions_list).to eq([
"Whisk everything into a smooth batter.",
"Fry in a hot pan until golden on both sides."
])
end
endA full spec covers every field, not only these three. Any spec under spec/sites/ shows the set:
the parsed ingredients, the groups, the metadata, the nutrients and their parsed form, and one
link. A field the page does not publish is asserted as nil. For a recipe with groups, assert each
purpose and how many lines it holds:
it "splits the ingredients into the groups the page names" do
expect(recipe.ingredient_groups.map { |group| [group.purpose, group.ingredients.size] }).
to eq([["COATING", 2], ["CHURROS", 8]])
endCheck every expected value against the page in a browser. A value copied from the gem's output without that check proves nothing.
bundle exec rspec spec/sites/com/example_spec.rbThe first run fetches the page and records it in spec/cassettes/com/example.yml. Every later run
replays the recording. See Testing. Try a few other recipes of the site in the console
as well, to catch the cases one page does not show.
Add the host to Supported Sites, in alphabetical order, as
- [example.com](https://example.com/). A site registered with also: gets one line per host.
Raise the number of sites at the top of the page by the lines you added. It matches
RecipeScrapers::Registry.hosts.size.
Commit the site file, the spec, the cassette and the updated list together, run bin/ci, and
open a pull request.
Using the gem
Adding a site
Copyright