Can I baseline CoPilot against other tools without committing. #205651
Replies: 8 comments 4 replies
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
This comment was marked as spam.
|
Hi @joelbrayman-collab — I regularly run read-only multi-tool benchmarks like this on my own repos (baselining Claude Code / Codex / Copilot against each other on architecture-reconstruction tasks), so this is a familiar setup. Short answers to your three questions: 1. Can you enable the policy with a free entitlement? No. That error comes from the Copilot desktop app's plan/interactive session features, which are org/enterprise-policy-gated capabilities. There is no personal-account setting that unlocks them — the toggle lives on the organization side (org owners on business/enterprise plans). You're not missing a checkbox; the desktop session feature simply isn't available on individual/free entitlements. (Known limitation — see discussion #205473 for the identical symptom.) 2. What's the minimum plan? For the desktop app's policy-gated sessions: that's an org/enterprise-plan decision, not something a free individual account can upgrade into. But you don't need that path at all — The route that works today without upgrading: use the Copilot extension inside VS Code attached to your existing local clone. Copilot's IDE agent mode does not require the org policy that the desktop app asks for; individual plans get a monthly agent-request quota (small, but ample for a benchmark run — check the current Copilot pricing docs for the exact number). To keep the run genuinely read-only and safe, benchmark against a fresh clone rather than your main working copy; that also protects against any agent that tries to "help" by writing files. 3. Trial? I wouldn't count on a benchmark-friendly trial for the higher individual tiers; current tiers change quickly, so check the pricing/compare page when you're ready to upgrade. Three things that decide whether this comparison is actually fair (learned the hard way benchmarking model gateways locally):
Your reconstruction-task design is a genuinely good benchmark: it can't be memorized from training data, so it measures real capability on your repo. Happy to share my scoring-rubric template if useful. If this answers your question, marking it as accepted helps other benchmarkers find it — cheers! 🙌 |
|
Thanks — this is extremely helpful and exactly the clarification I was looking for. I will test through the VS Code Copilot extension against a fresh clone and record the exact model ID, plan/date and usage. I agree that multiple runs are necessary.
Yes, please share your scoring-rubric template. I would like to use a consistent rubric across Copilot, Codex and Claude Code.
From: Lx 🎀 ***@***.***>
Date: Sunday, August 23, 2026 at 13:30
To: community/community ***@***.***>
Cc: joelbrayman-collab ***@***.***>; Mention ***@***.***>
Subject: Re: [community/community] Can I baseline CoPilot against other tools without committing. (Discussion #205651)
Hi @joelbrayman-collab<https://github.com/joelbrayman-collab> — I regularly run read-only multi-tool benchmarks like this on my own repos (baselining Claude Code / Codex / Copilot against each other on architecture-reconstruction tasks), so this is a familiar setup. Short answers to your three questions:
1. Can you enable the policy with a free entitlement? No. That error comes from the Copilot desktop app's plan/interactive session features, which are org/enterprise-policy-gated capabilities. There is no personal-account setting that unlocks them — the toggle lives on the organization side (org owners on business/enterprise plans). You're not missing a checkbox; the desktop session feature simply isn't available on individual/free entitlements. (Known limitation — see discussion #205473<#205473> for the identical symptom.)
2. What's the minimum plan? For the desktop app's policy-gated sessions: that's an org/enterprise-plan decision, not something a free individual account can upgrade into. But you don't need that path at all —
The route that works today without upgrading: use the Copilot extension inside VS Code attached to your existing local clone. Copilot's IDE agent mode does not require the org policy that the desktop app asks for; individual plans get a monthly agent-request quota (small, but ample for a benchmark run — check the current Copilot pricing docs for the exact number). To keep the run genuinely read-only and safe, benchmark against a fresh clone rather than your main working copy; that also protects against any agent that tries to "help" by writing files.
3. Trial? I wouldn't count on a benchmark-friendly trial for the higher individual tiers; current tiers change quickly, so check the pricing/compare page when you're ready to upgrade.
Three things that decide whether this comparison is actually fair (learned the hard way benchmarking model gateways locally):
1. Record the actual model ID, not the UI label. "Copilot" / "Claude Code" / "Codex" are product labels; each routes to a different underlying model depending on your plan and settings. Two tools that look like a fair fight are often not running comparable models. Put model ID + date + plan tier in your results table, or the comparison silently expires.
2. One identical prompt, and require analysis-only. Same prompt text for all three, and explicitly instruct each tool to reconstruct and report — no edits proposed, none applied. A tool that starts modifying the repo mid-run both breaks the read-only premise and contaminates its own score.
3. Pre-define a rubric and run each tool 2–3 times. Score each reconstruction area (git state, architecture, governance, deployment state, priorities) 0–3 against ground truth. Single-run variance across runs is frequently larger than the difference between tools, so a one-shot run per tool tells you almost nothing.
Your reconstruction-task design is a genuinely good benchmark: it can't be memorized from training data, so it measures real capability on your repo. Happy to share my scoring-rubric template if useful. If this answers your question, marking it as accepted helps other benchmarkers find it — cheers! 🙌
—
Reply to this email directly, view it on GitHub<#205651?email_source=notifications&email_token=CGWWRYYT5F4HK3VX6N6NW6L5LMS2JA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBRGI2TSNZYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18125978>, or unsubscribe<https://github.com/notifications/unsubscribe-auth/CGWWRY33BH4YAXL2NLNCJD35LMS2JAVCNFSNUABIKJSXA33TNF2G64TZHMZTAMJVG4ZTGNBUHNCGS43DOVZXG2LPNY5TCMBWG4YTCMRUUF3AE>.
Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS<https://github.com/notifications/mobile/ios/CGWWRY7UF322YSEAEI6N4F35LMS2JA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBRGI2TSNZYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSVGM33PORSXEX3JN5ZQ> and Android<https://github.com/notifications/mobile/android/CGWWRYZQLLWPXVLORCF4GJT5LMS2JA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBRGI2TSNZYUZZGKYLTN5XKO3LFNZ2GS33OUVSXMZLOOSXGM33PORSXEX3BNZSHE33JMQ>. Download it today!
You are receiving this because you were mentioned.
|
|
Hey, your hypothesis is basically confirmed by GitHub's own docs. "Copilot in GitHub Desktop" is explicitly listed as one of the org-level policy toggles (currently in public preview), separate from Enterprise-only gating. So the error you're hitting is almost certainly: your org owns the repo -> org's Copilot policy currently has that feature disabled/unset -> your personal entitlement can't override it, not "you need Enterprise." To your specific questions:
Bottom line: check with whoever owns your org's GitHub settings and ask them to flip the "Copilot in GitHub Desktop" org policy on, that's very likely the actual blocker, not your plan tier. |
Uh oh!
There was an error while loading. Please reload this page.
🏷️ Discussion Type
Question
💬 Feature/Topic Area
Copilot in GitHub
Body
Can I benchmark GitHub Copilot as a repository architecture agent before upgrading?
I want to evaluate GitHub Copilot against Claude Code and OpenAI Codex as an architecture/governance agent for a large existing private repository.
The repository is already hosted in GitHub and cloned locally. I have successfully benchmarked Claude Code and Codex by giving each read-only access to the repository and asking them to independently reconstruct Git state, architecture, governance, implementation state, deployment state, open/closed work, and recommended next priority.
I want to run the same read-only benchmark with GitHub Copilot before purchasing a higher plan.
I installed the GitHub Copilot desktop app and successfully attached the existing repository. However, both Plan and Interactive sessions fail immediately with:
“You are not authorized to use this Copilot feature, it requires an enterprise or organization policy to be enabled.”
The repository is owned by my GitHub organization.
I need to know:
Can I enable the required organization policy and run this benchmark using my current/free Copilot entitlement without purchasing Business/Enterprise?
If not, what is the minimum plan required: Pro, Pro+, Max, Business, or Enterprise?
Is there a trial/evaluation that will let me test the repository agent before committing to a paid plan?
I am specifically testing GitHub Copilot itself, not Claude Code or Codex as third-party agents.
Does the Copilot desktop agent have any rolling/session usage limit comparable to Claude's session limit or Codex's five-hour window, or is usage controlled only by the monthly AI-credit pool?
Can additional usage be given a hard spending cap with no automatic overage?
My objective is to determine whether Copilot can serve as a sustained repository-grounded software Architect/Governor, with Cursor remaining the separate implementation executor.
There is also a potentially important clue in GitHub's documentation that may explain the error without requiring Enterprise.
GitHub says organization/enterprise policies can control access to Copilot features, including agents and CLI capabilities. And a recent Community discussion specifically notes that an organization policy can take precedence even when someone is bringing their own individual Copilot entitlement, and that enabling the relevant policy does not itself necessarily mean purchasing organization Copilot licenses.
So the error we're seeing may be:
your organization owns Discovery-Search → organization policy currently blocks the desktop/agent capability → your personal Copilot entitlement cannot override that policy.
That is quite different from:
“You need GitHub Enterprise.”
And there is another official GitHub document worth noting: for an individual subscriber, GitHub says Copilot coding agent is available with Pro+ and can be enabled for selected repositories through the individual's Copilot settings.
So I would not purchase anything yet.
All reactions