Releases: Assessment-Nerd/codex-involution-guard
Release list
v0.2.1 — Evidence reuse and budgeted evaluation
English
Changes
- Reuse trustworthy verification evidence matching the current requirement, artifact version and relevant environment. Check only missing or invalidated evidence; necessary safety checks remain.
- Added an optional, budgeted restricted-action evaluation script and sanitized results. It is not a runtime dependency or a full Codex benchmark.
- English-first bilingual README now focuses on current use. Version history lives in Releases; public MIT project contributions are welcome through Issues and PRs.
Evidence and limits
24 episodes / 40 API calls: no skill 3/8, v0.2 3/8, candidate 5/8 success. Redundant checks 1/2/1 respectively. Total check counts did not decline versus v0.2; candidate API cost increased about 1.9%.
Estimated uncached-rate API cost: $0.01189395, not an invoice or total development cost. Small samples do not establish general gains, safety or savings. Three candidate failures remain. No new paid run for release packaging.
Method and failures · Installation · Feedback
Experimental prerelease. The tag includes the documentation/release organization update; core skill behavior is unchanged from the tested correction.
한국어
변경 사항
- 현재 요구·아티팩트 버전·관련 환경에 맞는 신뢰할 만한 검사 결과는 재사용합니다. 빠지거나 무효화된 증거만 다시 확인하며 필수 안전 검증은 유지합니다.
- 선택적인 예산 제한 행동 평가 스크립트와 합성 결과를 추가했습니다. 실행 의존성이나 Codex 전체 벤치마크는 아닙니다.
- README는 영어 먼저·한국어 다음으로 현재 사용법을 안내합니다. 버전 이력은 Releases에서 관리하며 공개 MIT 프로젝트로 Issues와 PR 기여를 받습니다.
검증과 한계
24회 실험·40 API 호출에서 스킬 없음 3/8, v0.2 3/8, 수정안 5/8 성공했습니다. 중복 검사는 각각 1/2/1회입니다. 전체 검사 횟수는 v0.2 대비 줄지 않았고 수정안 비용은 약 1.9% 늘었습니다.
공식 단가 기준 추정 API 비용은 $0.01189395이며 청구 확정액이나 총개발비가 아닙니다. 작은 표본으로 일반 성능·안전·절감을 입증하지 못했고 수정안 실패 3건도 남았습니다. 릴리스 정리를 위한 추가 유료 실험은 없습니다.
위 링크에서 설치·평가·실패 기록과 피드백 경로를 확인할 수 있습니다. 실험적 사전 릴리스이며 태그에는 문서 정리도 포함하지만 스킬 동작은 평가한 수정안과 같습니다.
v0.2.0 — Behavioral-science-informed guidance
English
Clarified three decision rules using behavioral-science-informed design: consider subtraction/reuse, weigh future value rather than sunk effort, and distinguish outcome progress from useful information. Added philosophy-to-implementation documentation. Human research is design inspiration, not proven LLM mechanisms.
Six fixed cases passed in both versions; no incremental behavioral benefit established. A potentially redundant recheck was recorded. No new runtime dependency.
Retrospective historical release at the v0.2 implementation commit.
한국어
행동과학을 참고해 재사용·통합·생략 대안, 매몰비용 대신 미래 가치 판단, 결과 진전과 유용한 정보의 구분을 명확히 했습니다. 철학에서 구현까지 설명을 추가했습니다. 인간 연구는 설계 근거이지 LLM에서도 입증된 기제라는 뜻이 아닙니다.
고정 사례 6개는 전후 모두 통과했습니다. 추가 효과는 미입증이며 불필요할 수 있는 재검사 제안도 기록했습니다. 실행 의존성은 추가하지 않았습니다.
v0.2 구현 커밋을 가리키는 과거 릴리스를 지금 등록합니다.
v0.1.0 — Initial experimental skill
English
Initial experimental instruction skill: preserve original outcomes, reuse existing work, avoid ineffective retries and distinguish evidence levels. No required server or API key. Initial six-case pilot met criteria with and without the skill; no incremental benefit demonstrated.
This historical release is being recorded retrospectively at the original implementation commit. It is not a claim that a GitHub Release existed at that time. MIT-licensed; feedback and pull requests welcome.
한국어
최초 실험적 지침형 스킬입니다. 원래 목표 보존, 기존 작업 재사용, 무의미한 반복 점검, 증거 수준 구분을 제공합니다. 서버·API 키가 필요 없습니다. 초기 6개 사례는 사용·미사용 모두 통과해 추가 효과는 입증되지 않았습니다.
원래 구현 커밋을 가리키는 과거 버전 기록을 지금 등록합니다. 당시 GitHub Release가 있었다는 뜻은 아닙니다. MIT 라이선스이며 피드백과 PR을 환영합니다.