You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Note on robots.txt: the GEO scanner's robots_txt check reports found: false for the docs site because it probes the domain root (https://github.github.com/robots.txt), which doesn't include the /gh-aw/ GitHub Pages project-site base path. A direct, authoritative check of https://github.github.com/gh-aw/robots.txt confirms it does exist (HTTP 200) and explicitly allows all major AI crawlers (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, anthropic-ai, Google-Extended, etc.) plus a Sitemap: directive. This is a scanner limitation, not a real site gap — no action needed on robots.txt for the docs site.
✅ Top Strengths
Meta tags: Docs site scores 14/14 — title, description, canonical URL, and full OpenGraph tags all present.
Schema JSON-LD: Docs site includes WebSite, Organization, SoftwareApplication, and FAQPage types (13/16).
Content structure: README scores 12/12 on content (word count, headings, links); docs site has a clear H1, 49 headings, and heading hierarchy.
Signals: Docs site has lang="en", an RSS feed (/blog/rss.xml), and a freshness date — full 6/6.
Brand & entity: Docs site links to Wikipedia, Wikidata, LinkedIn, and Crunchbase for Knowledge Graph disambiguation (7/10).
Multimodal: Alt-text coverage is 100% and video captions are present (readiness: "good").
CDN/bot access: No AI bots (GPTBot, PerplexityBot, Claude-SearchBot, Googlebot, etc.) are blocked or challenged at the CDN level.
README: robots.txt (15/18) and llms.txt (14/18) checks both pass strongly at the repo level.
🚨 Critical Gaps
No llms.txt on the docs site (0/18) — no /llms.txt file exists to help AI systems index the site structure.
No AI Discovery signals on either the docs site or README (0/6 each) — missing /.well-known/ai.txt, /ai/summary.json, /ai/faq.json, /ai/service.json.
README has no Schema JSON-LD (0/16) — no WebSite/Organization/FAQPage structured data on the GitHub repo homepage.
Hidden text detected on both docs site and README (display:none/visibility:hidden with content) — flagged as a potential cloaking pattern that AI crawlers may penalize (contributes to the -3 negative penalty on both audits).
Keyword stuffing: "github" density is 2.7% (docs) / 3.6% (README) — vocabulary diversification recommended.
4 of 20 sampled docs pages score "critical" (worst: /blog/9 at 33, /blog/11 at 35, /blog/8 at 35, /about at 35) — largely driven by weak schema/brand-entity scores on those pages.
🔧 Recommended Fixes
Ordered by estimated impact:
Create /llms.txt for the docs site (up to +18 pts) — run geo llms --base-url https://github.github.com/gh-aw to generate a structured AI-indexing file with H1, description, sections, and links.
Add Organization + WebSite JSON-LD schema to the README (up to +16 pts) — include name, url, logo, sameAs (Wikipedia/Wikidata/LinkedIn/Crunchbase), and FAQPage schema.
Add AI-discovery well-known files to the docs site and README (up to +6 pts each): /.well-known/ai.txt, /ai/summary.json, /ai/faq.json, /ai/service.json.
Remove or reveal hidden text (display:none/visibility:hidden elements with real content) on both the docs site and README to eliminate the cloaking-pattern penalty (-3 pts each).
Front-load key information in the first 30% of docs-site content for better AI snippet selection.
Add numerical data/concrete statistics to docs content (~+40% AI visibility per GEO heuristics).
Diversify vocabulary to reduce "github" keyword density below stuffing thresholds.
Add <link rel="canonical"> to the README page to prevent duplicate-content indexing issues.
Investigate the 4 critical-band docs pages (/blog/9, /blog/11, /blog/8, /about) for schema and brand-entity gaps specifically.
📋 Full Breakdown by Category
Docs site (https://github.github.com/gh-aw/) — score 44/100:
Category
Score
Max
Robots.txt
0*
18
llms.txt
0
18
Schema JSON-LD
13
16
Meta Tags
14
14
Content
7
12
Signals
6
6
AI Discovery
0
6
Brand & Entity
7
10
Negative penalty
-3
—
*See robots.txt note above — this is a scanner domain-root artifact; the actual /gh-aw/robots.txt is valid and AI-bot-friendly.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
GEO Audit Report — github/gh-aw
Audit Date: 2026-08-18
Run: https://github.com/github/gh-aw/actions/runs/32160801176
📊 Scores
github.github.com/gh-aw/)Note on robots.txt: the GEO scanner's
robots_txtcheck reportsfound: falsefor the docs site because it probes the domain root (https://github.github.com/robots.txt), which doesn't include the/gh-aw/GitHub Pages project-site base path. A direct, authoritative check ofhttps://github.github.com/gh-aw/robots.txtconfirms it does exist (HTTP 200) and explicitly allows all major AI crawlers (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, anthropic-ai, Google-Extended, etc.) plus aSitemap:directive. This is a scanner limitation, not a real site gap — no action needed on robots.txt for the docs site.✅ Top Strengths
WebSite,Organization,SoftwareApplication, andFAQPagetypes (13/16).lang="en", an RSS feed (/blog/rss.xml), and a freshness date — full 6/6.🚨 Critical Gaps
llms.txton the docs site (0/18) — no/llms.txtfile exists to help AI systems index the site structure./.well-known/ai.txt,/ai/summary.json,/ai/faq.json,/ai/service.json.WebSite/Organization/FAQPagestructured data on the GitHub repo homepage.display:none/visibility:hiddenwith content) — flagged as a potential cloaking pattern that AI crawlers may penalize (contributes to the -3 negative penalty on both audits)./blog/9at 33,/blog/11at 35,/blog/8at 35,/aboutat 35) — largely driven by weak schema/brand-entity scores on those pages.🔧 Recommended Fixes
Ordered by estimated impact:
/llms.txtfor the docs site (up to +18 pts) — rungeo llms --base-url https://github.github.com/gh-awto generate a structured AI-indexing file with H1, description, sections, and links.name,url,logo,sameAs(Wikipedia/Wikidata/LinkedIn/Crunchbase), andFAQPageschema./.well-known/ai.txt,/ai/summary.json,/ai/faq.json,/ai/service.json.display:none/visibility:hiddenelements with real content) on both the docs site and README to eliminate the cloaking-pattern penalty (-3 pts each).<link rel="canonical">to the README page to prevent duplicate-content indexing issues./blog/9,/blog/11,/blog/8,/about) for schema and brand-entity gaps specifically.📋 Full Breakdown by Category
Docs site (
https://github.github.com/gh-aw/) — score 44/100:*See robots.txt note above — this is a scanner domain-root artifact; the actual
/gh-aw/robots.txtis valid and AI-bot-friendly.README (
https://github.com/github/gh-aw) — score 55/100:Sitemap-wide averages (20 pages sampled of 224 discovered):
Band distribution: 16 pages "foundation", 4 pages "critical" (out of 20 audited).
📄 Sitemap Page Scores
Top pages:
/gh-aw(home)/blog/2026-01-13-meet-the-workflows-continuous-improvement/blog/2026-01-13-meet-the-workflows-continuous-refactoring/blog/2026-01-12-welcome-to-pelis-agent-factory/blog/2026-01-13-meet-the-workflows-advanced-analyticsWorst pages:
/blog/9/blog/11/blog/8/about/blog/10Automated audit powered by geo-optimizer-skill · Run logs
All reactions