Skip to content

refactor(skills): simplify current guidance and preserve scan evidence - #98

Merged
drewstone merged 1 commit into
mainfrom
refactor/skills-current-guidance-20260905
Sep 5, 2026
Merged

refactor(skills): simplify current guidance and preserve scan evidence#98
drewstone merged 1 commit into
mainfrom
refactor/skills-current-guidance-20260905

Conversation

@drewstone

Copy link
Copy Markdown
Owner

Skills repeated full workflows, prescribed arbitrary staffing and measurement thresholds, and copied changing API and model facts.
The revised guidance keeps essential decisions in each entrypoint and loads named supporting procedures only when their conditions apply.
Changing facts point to their current owning source.
Existing consumers still check the package and execution path they actually use.

This covers all 61 dotfiles-owned skills: 59 entrypoints changed; site-clone and session-continuity remain byte-for-byte unchanged.
The related Runtime skill update is tangle-network/agent-runtime#1102.

  • Delete 26 competing full-reference documents, the duplicate skill ladder, three redundant shell helpers, and copied scan/loop workflows.
  • Remove fixed worker, edit, repetition, score, report, and automation quotas while preserving required behavior and evidence.
  • Correct model defaults, research selection criteria, API ownership, release completion checks, and append-only history guidance.
  • Preserve security boundaries, durable recovery, experiment identities, comparison integrity, and conditional skill chaining.
  • Allow a checked no-change result and short evidence-backed status answers; align the shared defaults with that behavior.
  • Fix Semgrep merging to preserve complete runs, zero-finding scans, invocation status, rule indexes, paths, and distinct findings.
    Reject malformed or unsupported input without publishing a partial merge; keep issue deduplication in triage.

Validation:

  • RTK_GUARD_MEASUREMENT_CACHE=/tmp/rtk-measurement-aliases-20260904.json npm test: 58 passed, 0 failed, 7 Linux-only skips; smoke passed.
    The unchanged RTK measurement cache was reused; its guard and mutation tests ran.
  • 61/61 skill frontmatter validations; final local-reference, logging, and chaining checks passed.
  • 7/7 Semgrep behavior regressions passed, including zero findings and preservation of existing output on invalid input.
    These are combiner checks, not complete SARIF schema validation.
  • 34 external source links checked against current owning files, trees, or repositories.
  • Independent task checks: report identified an incomplete release despite 14 passing tests and HTTP 200; simplify preserved all four files in a sound package.
  • Executed the finalize reconstruction, scan filtering, installer preservation, and docs-scanner examples with positive and relevant negative cases.
  • Independent source review found one shared reporting contradiction; it was corrected and re-reviewed with no remaining findings.

These checks support instruction integrity and the tested behaviors.
They do not establish a comparative improvement in agent productivity.
Unchanged bundled lint rules were not tested against a fresh Oxlint installation.

Text measurements use whitespace-separated words and physical lines across all 61 owned entrypoints and their supporting Markdown.
Shared author guidance is excluded from these totals.
Baseline is 2178ec6; candidate is 739a096.

Measure Before After
Entrypoint words 39,805 25,234
Entrypoint lines 4,483 3,332
Supporting Markdown files 42 39
Supporting Markdown words 42,460 11,426
Entrypoint words: min / median / p90 / max 107 / 550 / 1,318 / 1,671 107 / 410 / 525 / 622
Complete per-skill measurements (before → after)
Skill Entrypoint words Entrypoint lines Supporting files Supporting words
agent-behavior-audit 330 → 356 42 → 45 1 → 1 797 → 351
agent-eval 483 → 337 58 → 51 0 → 0 0 → 0
agent-integrations-adoption 532 → 364 72 → 53 0 → 0 0 → 0
arena-experiment 462 → 460 82 → 82 0 → 0 0 → 0
autopsy 377 → 373 44 → 46 1 → 1 1181 → 403
breakout 770 → 464 66 → 63 1 → 1 912 → 341
build-agent-app 1364 → 494 188 → 66 0 → 2 0 → 393
calibrate-before-measure 386 → 428 50 → 53 0 → 0 0 → 0
converge 341 → 364 51 → 46 1 → 0 1291 → 0
critical-audit 1318 → 572 119 → 68 6 → 5 3636 → 628
deep-clean 1373 → 454 113 → 56 1 → 0 692 → 0
deploy-proof 756 → 432 96 → 51 0 → 0 0 → 0
diagnose 1290 → 499 120 → 66 1 → 0 512 → 0
director-autopsy 404 → 431 59 → 59 0 → 0 0 → 0
discovery-lead 520 → 468 62 → 58 0 → 1 0 → 290
docs-slop-audit 298 → 335 44 → 48 2 → 1 1146 → 465
dont-collapse-the-architecture 295 → 317 42 → 44 0 → 0 0 → 0
eval-agent 677 → 486 101 → 65 0 → 1 0 → 338
eval-engineering 882 → 558 123 → 72 0 → 1 0 → 255
eval-harness-diagnose 679 → 503 92 → 64 0 → 0 0 → 0
evolve 1454 → 593 121 → 72 4 → 3 5770 → 1393
finalize 869 → 483 60 → 57 1 → 1 2503 → 724
governor 1400 → 622 111 → 73 0 → 0 0 → 0
ground-truth 586 → 490 38 → 59 0 → 0 0 → 0
harden 377 → 399 51 → 49 1 → 1 1852 → 484
harness-escalation-audit 564 → 430 79 → 55 0 → 0 0 → 0
hub-sdk-adoption 619 → 369 82 → 53 0 → 0 0 → 0
hypothesize 829 → 419 59 → 57 1 → 1 863 → 233
install-anti-slop 555 → 341 97 → 48 0 → 1 0 → 108
meta-harness 361 → 460 51 → 59 1 → 1 2024 → 420
model-freshness 598 → 380 86 → 55 0 → 0 0 → 0
multi-pursue 353 → 297 45 → 45 1 → 0 832 → 0
nano-banana 309 → 283 78 → 41 0 → 0 0 → 0
operate 544 → 468 73 → 63 0 → 0 0 → 0
orchestrate 550 → 327 66 → 48 1 → 1 492 → 200
polish 1145 → 386 103 → 51 1 → 1 420 → 253
problem-sourcing 523 → 314 43 → 45 0 → 1 0 → 199
product-design-audit 354 → 410 48 → 51 2 → 1 2797 → 408
product-design 369 → 380 50 → 51 1 → 1 1308 → 242
product-innovation-audit 371 → 429 56 → 50 1 → 1 3111 → 343
pursue 1445 → 594 121 → 70 1 → 0 758 → 0
push-past-easy 486 → 349 43 → 44 0 → 0 0 → 0
reconcile 1006 → 563 97 → 65 0 → 0 0 → 0
reflect 1070 → 525 114 → 65 1 → 2 1160 → 589
refresh-reasoning-capabilities 736 → 316 109 → 46 0 → 0 0 → 0
release-conductor 338 → 384 50 → 48 1 → 1 1061 → 297
report 1671 → 447 148 → 55 0 → 1 0 → 515
review-to-green 688 → 440 57 → 54 0 → 0 0 → 0
sandbox-sdk-integration 335 → 329 46 → 51 1 → 2 469 → 394
semgrep 316 → 444 56 → 62 4 → 3 2360 → 605
session-continuity 359 → 359 53 → 53 0 → 0 0 → 0
ship 315 → 344 50 → 45 1 → 0 889 → 0
signal-distill 292 → 362 44 → 51 1 → 0 868 → 0
simplify 581 → 392 70 → 51 1 → 0 1061 → 0
site-clone 107 → 107 19 → 19 0 → 0 0 → 0
slack-alerts 1057 → 353 72 → 50 0 → 1 0 → 205
tangle-blockchain-blueprint 668 → 396 86 → 55 0 → 0 0 → 0
tangle-ops 874 → 445 71 → 60 0 → 0 0 → 0
ui-test 293 → 336 45 → 44 1 → 1 875 → 350
verify 637 → 416 69 → 54 0 → 0 0 → 0
writing-profile 264 → 358 42 → 52 1 → 0 820 → 0

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant