Skip to content

Sources

Mazhar Ahmed edited this page Aug 28, 2026 · 4 revisions

Sources

A source is one collection QQL can resolve against. Each has a short code, and the code decides how the numbers in a reference are read.

Built in

Code Collection primary is Range CODE::N
Q Quran Surah 1–114 1–6236
B Sahih al-Bukhari chapter (kitab) 1–97 1–7563
M Sahih Muslim chapter 1–56 1–7563
AD Sunan Abi Dawud chapter 1–43 1–5274
T Jami' at-Tirmidhi chapter 1–49 1–3956
N Sunan an-Nasa'i chapter 1–51 1–5758
IM Sunan Ibn Majah chapter 1–37 1–4341
MA Muwatta Malik chapter 1–61 1–1858
DA Sunan ad-Darimi † chapter 1–23 —
RS Riyad as-Salihin chapter 1–19 —
BM Bulugh al-Maram chapter 1–16 —
AM Al-Adab Al-Mufrad chapter 1–57 —
MK Mishkat al-Masabih chapter 1–24 —
SM Ash-Shama'il Al-Muhammadiyah chapter 1–56 —
NW Al-Arba'in an-Nawawiyyah chapter 1 1–42
QD Forty Hadith Qudsi chapter 1 1–40
SW Forty Hadith of Shah Waliullah chapter 1 1–40
HM Hisnul Muslim (alias HISN) chapter 1–132 1–267

Codes are case-insensitive. qql --sources prints the registered list, including any custom sources.

The three forties are undivided books, so their whole text is chapter 1 and NW::13 and NW:1:13 name the same hadith.

† ad-Darimi is Arabic only. Upstream carries no English for any of its 2,757 hadiths, so every DA record has an empty en. Everything works — the records resolve, the indexes are built — but an English search finds nothing there, and only there. Every other collection is 99–100% translated.

Six collections have no book-wide numbering

DA, RS, BM, AM, MK and SM carry a dash in the last column. That is not an omission — QQL has no citation numbering it can source for them. The upstream data gives a position within the chapter and nothing else, and the sequential position within the book is not the number these works are cited by.

Publishing it as one would answer real citations with the wrong hadith, which is the same trap idInBook sets for Bukhari (see Architecture). So the flat form is refused outright:

DA::1        QQL_UNSUPPORTED — "DA has no canonical citation numbering;
                               address it as DA:chapter:number"
DA:1:1       fine

The refusal reaches exactly two forms — CODE::N, and the unscoped exact search CODE:"…", which would have to walk that same axis. Scoped search and both indexed engines never needed a book-wide number and are unaffected:

DA:"رسول"      refused
DA:1:"رسول"    118 results
DA:?"رسول"     ranked
DA:*"رسول"     ranked

(Arabic here because ad-Darimi has no English — see the † note above. For any other chaptered collection the same four forms work with English.)

Musnad Ahmad ibn Hanbal is not carried at all: upstream has 8 of its musnads and 1,374 of roughly 27,000 hadiths — too incomplete to publish under a code that implies the whole collection.

Two ways to address an item

Within a chapter — SOURCE:primary:n counts from 1 inside primary:

Q:2:255      Surah 2, ayah 255
B:1:1        the first hadith of chapter 1
HM:27:1      the first supplication of chapter 27

Across the whole book — SOURCE::n, the canonical citation numbers ('Abd al-Baqi for Bukhari, Dar-us-Salam for Muslim, sunnah.com reference numbers for the rest):

Q::100       the 100th ayah of the mushaf   (= 2:93)
B::6403      what every hadith site cites as Bukhari 6403
HM::75       the 75th supplication          (in chapter 27)

Bounds are in the table above. The hadith numbers resolve through committed maps (sources/canonical/*.json) validated against the text; front matter and lettered variants are holes — an error alone, skipped inside a range. Ten collections have such a map; the six listed below the table do not, and refuse the form.

Records from the flat form carry "numbering": "book" so the two schemes can never be confused in a mixed response.

What a record looks like

Every record carries source, collection, ar and en. The rest is source-specific rather than forced into one shape.

Quran

{ "source": "Q", "collection": "Quran", "surah": 2, "surah_name_ar": "البقرة",
  "surah_name_en": "Al-Baqarah", "ayah": 255, "ar": "…", "en": "…" }

Hadith

{ "source": "B", "collection": "Sahih al-Bukhari", "chapter": 1,
  "chapter_name_ar": "كتاب بدء الوحى", "chapter_name_en": "Revelation",
  "number": 1, "narrator": "Narrated 'Umar bin Al-Khattab:", "ar": "…", "en": "…" }

Hisnul Muslim

{ "source": "HM", "collection": "Hisnul Muslim", "chapter": 27,
  "chapter_title": "Words of remembrance for morning and evening",
  "number": 1, "repeat": 1, "note": "…", "audio": "…", "ar": "…", "en": "…" }

repeat is how many times to say it. note is usually a transliteration but sometimes a recitation instruction — upstream stores both under one key, so it is exposed under a neutral name.

Where the text comes from

Read straight from sources/ in each project's own layout. There is no build-time transform and no second copy.

Source Path Notes
Quran sources/quran/chapters/{1..114}.json generated, committed — see below
Hadith sources/hadith/{book}/{chapter}.json committed, 16 collections from hadith-json (ISC)
Hisnul Muslim sources/hisnul-muslim/husn_en.json committed, one file

Files load on first use and stay cached for the life of the context, so Q:2:255 reads about 50 KB rather than the whole mushaf.

Why the Quran text is generated

sources/quran/ is built by scripts/build-quran.py and committed. The Arabic comes from Tanzil's Uthmani text, because the quran-json-arabic package spells three combining marks with codepoints that mean something else (U+0657 where an open fathatan belongs, and so on) — enough to make 2:286 read isru instead of isran in any font that follows Unicode. Names, translation and transliteration still come from that project.

Rebuilding needs it, but only then:

git clone --depth 1 https://github.com/asim/quran-json-arabic /tmp/qja
python3 scripts/build-quran.py --meta /tmp/qja/dist/chapters/en

Data quirks QQL absorbs

Upstream data is authoritative and never rewritten, but it is not uniform:

  • Hisnul Muslim chapters are stored out of order — array position 0 is chapter 27 — so they are looked up by id, not position.
  • That file also has a UTF-8 BOM, two objects with duplicate keys, and one misspelled field.
  • Sunan an-Nasa'i has a chapter numbered 35.2, so chapter ids are not always integers.
  • Some collections have an introduction.json beside the numbered chapters, and a few have a lettered one (35b.json in Nasa'i, 8b.json in the Shama'il). QQL addresses chapters by number, so those are unreachable — by a reference and by a search alike.
  • The three forties ship upstream as a single all.json rather than numbered chapters. Each is carried here as 1.json; renaming it is the only edit made to any upstream file.
  • Translation coverage is not uniform. Sunan ad-Darimi has no English at all; the Muwatta has 125 entries with an English text but no Arabic. Both fields are always present in a record — an untranslated one is an empty string, never a missing key or a fabricated translation.

Text handling

Arabic passes through byte-for-byte. No Unicode normalization: not tashkeel, not Quranic marks, not zero-width characters. Invalid UTF-8 is rejected rather than lossily replaced.

Search folds Arabic marks for comparison only — see Search.

Clone this wiki locally