Repository navigation
Glossary
This glossary defines terms spanning search-engine optimisation, answer-engine
optimisation, generative-engine optimisation, and OmniRank's own fix-safety vocabulary —
from AnswerBlock and speakable markup to llms.txt, JSON-LD, fixTier and blast radius.
Each entry stands alone as a short, quotable answer to "what does this term mean,"
matching the same answer-first format OmniRank scores every audited page against.
Terms are grouped by theme, not alphabetised, so related concepts sit near each other.
Each H3 heading is a stable anchor — link directly to Glossary#term-name from anywhere.
SEO is the practice of structuring a webpage's HTML and content so crawlers like
Googlebot and Bingbot can find, parse and rank it — canonical tags, correctly sized title
and description metadata, and valid structured data are all part of it. OmniRank's SEO
layer checks h1, canonical, title-length, description-length, og, hreflang,
image-dims, schema, five site-level cross-URL gates, four on-page accessibility gates
(image-alt, heading-order, link-text, lang, new in v0.4.0), and five
indexability-contradiction gates (also new in v0.4.0, see Contradictions) against a
page's real HTML.
AEO is the practice of shaping a page so an AI Overview, Copilot, or voice assistant can lift a short, factual answer directly from it rather than sending a user to click through. OmniRank's AEO layer checks for a script-aware AnswerBlock (40–60 words for Latin scripts, 80–200 characters for CJK scripts), at least three FAQ pairs, and speakable markup that resolves to real content on the page.
GEO is the practice of making a site's content machine-ingestible ground truth that a
generative AI system such as ChatGPT, Claude, Perplexity or Gemini can retrieve and cite
in a generated answer. OmniRank's GEO layer checks that llms.txt, llms-full.txt and
facts.json are published and reachable, and that robots.txt does not block named AI
crawlers.
SMM is publishing and maintaining a consistent brand presence across social platforms so
entity resolvers and human audiences both recognise the same organisation everywhere it
appears. OmniRank's config schema reserves smm and sameAs fields for this, but no SMM
skill has shipped yet — smm-content and smm-publish are both roadmap items, targeted
at v0.6 and v0.7 respectively.
An answer engine is a search surface that returns a direct answer instead of, or above, a
list of links — Google's AI Overviews, Microsoft Copilot, and voice assistants are all
examples. Answer engines favour pages with a short, self-contained, liftable passage near
the top of the content, which is exactly what OmniRank's answer-block gate checks for.
A generative engine is a large-language-model-based system — ChatGPT, Claude,
Perplexity, Gemini — that synthesises a response from retrieved content rather than
simply linking to it. Generative engines need machine-readable ground truth and an
explicit citation licence before they reliably quote a source, which is what llms.txt,
llms-full.txt and facts.json exist to provide.
AI Overviews is Google's AI-generated summary shown above traditional search results for
many queries, synthesising an answer from multiple sources rather than listing links
alone. A page is more likely to be drawn into one when it carries a short, factual,
liftable passage — the same shape OmniRank's AEO layer checks for with its answer-block
gate.
RAG is the technique behind most generative-engine answers: a system retrieves relevant
passages from an external source at query time and feeds them into a language model's
context before it generates a response, rather than relying solely on what the model
memorised during training. llms-full.txt and facts.json exist specifically to give a
RAG pipeline clean, structured passages to retrieve.
In generative-engine terms, a citation is a generated answer that names or links its
source rather than presenting synthesised text as if it had no origin. OmniRank cannot
make any engine cite a source — no third party controls that — but an explicit citation
licence in llms.txt removes one concrete barrier.
An AnswerBlock is a plain-prose HTML element (OmniRank's default selector is
.answer-block) that directly answers "what is this page about" in subject-verb-object
sentences, with no lists and no superlatives. Its target length depends on the page's
script — 40–60 words for Latin text by default, 80–200 characters for CJK scripts — set
by aeo.answerBlock (see Configuration-Reference#aeoanswerblock). It is the passage
an answer engine is most likely to lift verbatim.
Speakable is a schema.org property, nested under speakable.cssSelector inside JSON-LD,
that names the CSS selectors of the page elements suitable for a voice assistant to read
aloud. OmniRank's speakable gate resolves every declared selector against the live page
and fails if any selector matches no element.
llms.txt is a curated, Markdown-formatted index proposed by the community convention at
llmstxt.org: the site name, canonical URL, a linked list of every
page with a one-line summary, and a citation licence block. OmniRank's geo-artifacts
skill generates it, and the llms-txt gate checks that /llms.txt returns HTTP 200. It
is a proposal, not a ratified standard — see FAQ#is-llmstxt-a-real-established-standard.
llms-full.txt is the full-corpus companion to llms.txt: every page's title, URL,
description and AnswerBlock text, concatenated into one Markdown file. It is the file
most often broken in production by the OpenNext/CloudFront 403 trap.
facts.json is the structured, JSON-formatted counterpart to llms.txt. As of v0.2.1,
its license field is no longer defaulted to CC-BY-4.0 when unset — it states plainly
that no reuse licence is granted, matching geo.license's "grant nothing" default.
JSON-LD (JavaScript Object Notation for Linked Data) is a format for embedding structured
data inside a <script type="application/ld+json"> tag so crawlers can parse an
unambiguous, typed description of a page's content without scraping visible text.
schema.org is the shared vocabulary of types — Organization, Article, FAQPage,
Review, and hundreds more — that JSON-LD, Microdata and RDFa markup all draw from.
Structured data is any markup — most commonly JSON-LD — that describes a page's content in a fixed, machine-parseable vocabulary rather than free text.
@graph is a JSON-LD keyword that groups multiple typed nodes under one shared
@context inside a single script block, instead of repeating @context on every node.
A canonical URL is the one absolute address a page declares, via <link rel="canonical">, as its authoritative version. OmniRank's canonical gate fails if the
tag is missing or the URL is relative rather than absolute — and, as of v0.2.0, its
site-level canonical-cluster gate additionally catches a canonical that points at
another page which itself canonicalises elsewhere (a chain).
hreflang is an HTML attribute on <link rel="alternate"> tags that tells a crawler
which URL serves which language or regional variant of a page. As of v0.2.0, OmniRank
also checks reciprocity across the whole crawled set — see hreflang-reciprocity in
Audit-Skill.
x-default is the reserved hreflang value that names the fallback page shown to a
visitor whose language or region matches none of a page's other declared alternates.
Crawl budget is the finite number of pages a search engine's crawler will fetch from a given site within a given time window. Wasting it on dead sitemap URLs, redirect chains or duplicate content leaves fewer crawls available for pages that actually matter.
IndexNow is a protocol, backed by Bing and Yandex, that lets a site push a URL directly
to a search engine's indexing queue the moment it changes. OmniRank's config schema
reserves an indexnowKeyFile field for it under indexing, a roadmap skill (target v0.3).
A sitemap is an XML file, conventionally at /sitemap.xml, listing every URL a site
wants crawled along with an optional lastmod date. As of v0.2.1, a missing or
unreachable sitemap now emits seo.sitemap.missing (error) instead of silently auditing
just the homepage.
lastmod is the XML element inside a sitemap entry that states when a URL last changed.
A lastmod value re-stamped to today's date on every build is a false freshness signal —
what OmniRank's lastmod-inflation gate detects when more than 90% of sampled entries
share one date.
robots.txt is a plain-text file at a site's root that tells crawlers which paths they
may or may not fetch. OmniRank's ai-allowlist gate reads a site's real, published
robots.txt and fails if any of 19 named AI-crawler user agents are fully disallowed —
this is also a "protected surface" that no fix mode will ever write to unattended, see
Fix-Tiers-and-Applicability#what-are-protected-surfaces.
Crawl hygiene is the practice of keeping every URL a crawler might encounter resolving to
a live, correctly coded response. As of v0.2.1, omnirank audit automatically runs
check_sitemap() (a redirecting or dead sitemap entry); check_removed() still requires
calling directly with your own list of retired URLs.
410 Gone is an HTTP status code that tells a crawler a resource was intentionally and
permanently removed, as distinct from 404 Not Found.
Mixed content is an https:// page requesting a subresource over literal http://.
OmniRank splits it by what browsers actually do about it: active mixed content
(<script>, <iframe>, stylesheets) is blocked outright and reported as error;
passive mixed content (<img>, favicons) is silently upgraded to https:// first and
reported only as warning, since OmniRank cannot verify the upgrade succeeded from the
HTML alone. See Security-Layer.
HSTS is a response header (Strict-Transport-Security) that tells a browser to refuse
ever downgrading a site to plain http://, for a stated max-age. OmniRank's hsts gate
reports its absence or a short max-age as an info-severity fact — never graded,
because what counts as "long enough" is a threat-model judgement outside an SEO tool's
scope. See Security-Layer.
CSP is a response header or <meta http-equiv> tag that restricts which sources a page
may load scripts, styles and other resources from. OmniRank's csp gate reports only
total absence, and parses a present policy for exactly one thing —
upgrade-insecure-requests, which suppresses the mixed-content findings — never grading
the policy's contents. See Security-Layer.
A contradiction, in OmniRank's v0.4.0 terminology, is a defect provable purely from a
site's own declarations disagreeing with each other — a sitemap URL its own robots.txt
disallows, a canonical pointing at a noindexed page, an hreflang alternate that is
itself noindexed. Every contradiction finding is 100% precision by construction: OmniRank
never has to trust an external source to know the two declarations conflict. See
Contradictions.
fixTier is the epistemic axis of OmniRank's fix model: what kind of information the
correct edit requires, as a static property of a finding id — mechanical, templated,
drafted, advisory or infrastructure. It is declared for all 67 finding ids in
scripts/py/omnirank/registry.py and replaces the removed autoFixable field. See
Fix-Tiers-and-Applicability.
applicability is the safety axis of OmniRank's fix model: whether one particular
occurrence of a finding may be applied unattended, computed as the minimum of the tier
ceiling, the locator's confidence, the edit's blast radius, and any protected-surface
ceiling — safe, unsafe or display-only. Every input can demote; none can promote.
See Fix-Tiers-and-Applicability.
The locator is the OmniRank component that resolves a finding's URL to {path, line, confidence} in the source tree — the answer to "which file is wrong?" It never guesses:
an unresolvable route returns none confidence rather than a plausible-looking path. See
The-Locator.
Blast radius is how many routes a located source file serves. A file serving exactly one
route can safely carry a route-specific literal (like a canonical URL); a shared layout
serving thousands cannot — writing one there would collapse the whole site to one
indexed page. routes_served is one of the four ceilings applicability takes the
minimum of. See Fix-Tiers-and-Applicability.
A protected surface is one of four categories of edit — robots.txt/crawler directives,
noindex/sitemap membership, canonical/hreflang sets, or any licence grant — that is
hard-capped at unsafe regardless of tier or locator confidence, and never reachable by
an unattended write at any point on the roadmap. See
Fix-Tiers-and-Applicability#what-are-protected-surfaces.
An entity, in search and AI-retrieval terms, is a distinct, identifiable thing — a person, organisation, product or place — that a search engine or generative engine tries to resolve to one canonical representation rather than treating as free text.
sameAs is a schema.org property listing the URLs of a site's other verified profiles.
In OmniRank's config, a null value in the sameAs map is not "not applicable"; it is a
documented entity-linking gap, and only non-null values are ever emitted into
facts.json.
A knowledge graph is a search engine's internal database of entities and the relationships between them, used to answer factual queries directly.
E-E-A-T stands for Experience, Expertise, Authoritativeness and Trustworthiness — the criteria Google's Search Quality Rater Guidelines describe for evaluating content. It is a rating framework, not an algorithmic score OmniRank or any third-party tool computes directly.
- FAQ — direct answers to common questions, including several referenced above
- Audit-Skill — the gates that check many of these terms in practice
- Security-Layer / Contradictions — the security and contradiction terms above, in full
- Fix-Tiers-and-Applicability — the fix-safety terms defined above, explained in full
-
GEO-Artifacts-Skill —
llms.txt,llms-full.txtandfacts.jsonin full detail - Research-and-Evidence — the published research behind the GEO discipline
OmniRank · maintained by S M Moshiur Rahman at TICON System Limited, Dhaka · Code MIT, docs CC BY 4.0 · v0.4.0 ships two skills plus the fix diff preview; the rest is Roadmap
One page. Every engine.
Start here
Skills
Fix preview
Reference
Learn
Project