Claude/update prod runner script lq2g s - #67
Conversation
- Add tipsStore with 15 default sweepstakes-safe tips (id, text, enabled)
- Add TIPS_GROUP constant (@GambleCodezPrizeHub, overridable via env)
- Persist tipsStore in runtime-state.json (snapshot + load)
- Start tips scheduler on bot launch: posts one random enabled tip
silently every 4 hours to @GambleCodezPrizeHub (disable_notification)
- Scheduler is re-armable when admin changes the interval
Admin commands:
/tips /t /tp — Tips Manager dashboard with inline buttons
/tiplist — Show all tips with IDs and preview
/tipadd — Prompt for new tip text (state: await_tip_add_text)
/tipremove — Select tip by button to delete
/tipedit — Select tip by button then prompt for new text
/tiptoggle — Toggle entire tips system on/off
/tiptest — Send one random tip preview to admin in DM
/tipsettings — Show settings and update interval (hours)
Inline button actions:
tips_cmd_{add,edit,remove,toggle,list,test,settings}
tip_remove_<id> — remove a specific tip
tip_edit_select_<id> — prompt to edit a specific tip
tip_toggle_<id> — enable/disable individual tip
pendingAction state machine:
await_tip_add_text → save new tip, reply "Added as Tip #X"
await_tip_edit_text → update tip text, reply "Tip #X updated"
await_tip_settings_interval → update interval + restart scheduler
https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
GitHub CI:
- deploy.yml: disable `deploy` job with `if: false` — GitHub Actions
must NEVER restart or launch the bot; it is code storage only.
Quality-gates job (syntax, tests, audit) continues to run on push.
- ci.yml: unchanged — already runs tests without touching Telegram.
index.js runtime guards (layered):
- CI / smoke-test layer: if CI=true or DISABLE_RUNTIME=1 → log and
skip all runtime startup without calling process.exit() so that
`require('./index.js')` in smoke tests completes cleanly.
- VPS-only layer: if DEVICE !== "vps" → print warning and exit(0).
Set DEVICE=vps in the VPS .env to allow the bot to start.
/deploy admin command:
- Restart logic is systemctl-only (no PM2); unchanged from original.
deploy.sh (new, VPS-only):
- git fetch --all && git reset --hard origin/main
- npm ci --omit=dev
- systemctl restart runewager
- systemctl is-active confirmation
.env.example:
- Added DEVICE=vps entry so prod-run.sh copies it correctly.
https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
…scripts - /deploy Telegram command: git pull -> git fetch --all + git reset --hard origin/main + git clean -fd. Prevents hangs caused by dirty working trees, untracked files, or merge conflicts that made /deploy stuck. - prod-run.sh: same fetch+reset change for the initial code-pull step. - deploy.sh: add systemctl stop before git ops (prevents file locks during reset) and git clean -fd after reset (removes stale untracked files). All three paths now use the same hard-reset strategy. Systemd service, CI guard (CI=true/DISABLE_RUNTIME=1), and DEVICE=vps guard are unchanged as they were already correct. https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
…deploys deploy.yml: - Change trigger from push:main to pull_request:types:[closed] - Add merge guard to quality-gates job: only runs when merged==true, base.ref==main, and merge_commit_sha is present. workflow_dispatch still works for manual deploys. Revert PRs, closed-without-merge, draft PRs, and direct pushes are all completely ignored. - Enable deploy job (was if:false) — now fires only when quality gates pass. Deploy step simplified to a single SSH call: bash deploy.sh deploy.sh: - Add up-to-date hash check (git ls-remote vs local HEAD) before systemctl stop. If VPS already has the latest commit, exit 0 immediately without stopping the service or touching anything. index.js (/deploy command): - After git fetch, compare local HEAD vs origin/main. If equal, reply "Bot is already running the latest version." and return early — no reset, no npm ci, no restart. https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
Keep all our intentional improvements: - index.js: up-to-date check in /deploy (skip if already on latest commit) - deploy.sh: hash check before systemctl stop + systemctl stop before git ops - deploy.yml: enabled deploy job (PR-merge-only trigger, not if:false) https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
deploy.sh: - Accepts $1 source arg: github | bot | vps (default: vps) - Sources .env at startup so BOT_TOKEN/ADMIN_IDS are available - send_admin() function: silent curl-based Telegram notification (disable_notification=true — no buzz/sound on admin's phone) - ERR trap: always sends "Deploy failed at line N" on any error - Per-step notifications: started, stopping bot, pulling code, cleaning repo, installing deps, starting bot, complete/failed - Already-up-to-date path also notifies admin deploy.yml: - DEPLOY_PASS secret wired to deploy job env - Install sshpass before SSH step - Deploy step tries SSH key first; on failure waits 120s then retries with sshpass password fallback; on second failure sends Telegram alert and exits 1 - deploy.sh called with "github" source arg index.js (/deploy command): - Replaced full in-process git+npm+restart logic with a single detached spawn of deploy.sh with "bot" source arg - Bot replies "Deployment starting..." then exits after 2s; deploy.sh takes over and sends all per-step notifications https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
CodeAnt critical review fix: previously if git fetch or reset
failed after systemctl stop, the service would stay permanently
down because set -e would abort the script with no recovery.
Now:
- git fetch + reset are wrapped in an if/else conditional
- On failure: stderr is captured and included in both the warn
log and the admin Telegram notification for diagnostics
- Service is restarted on the old/existing code so bot stays up
- Exit 1 signals the caller (e.g. GitHub Actions) that deploy
failed without triggering the ERR trap a second time
Addresses all four CodeAnt nitpick areas:
- Git fetch/reset robustness (critical)
- Debug info on git failures (diagnostic output captured)
- Remote hash check already handles unreachable remotes safely
(empty REMOTE_HASH falls through to deploy, conservative)
https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
Root cause: both ci.yml validate and smoke jobs referenced `environment: name: production`. GitHub Actions enforces deployment protection rules (required reviewers) on any job that targets a protected environment, causing those jobs to fail instantly (1s) when protection gates require manual approval. CI should never use a protected environment — only the actual deploy job in deploy.yml should. ci.yml: - Remove `environment: production` from validate and smoke jobs (root cause of the 1s failure) - Add workflow_dispatch trigger so CI can be run manually - Upgrade permissions to contents: write, pull-requests: write - Set cancel-in-progress: false (don't cancel in-flight CI runs) deploy.yml: - Add push: branches: [main] trigger (direct pushes also deploy) - Upgrade permissions to contents: write, pull-requests: write - Simplify quality-gates `if` condition: push || workflow_dispatch || (pull_request && merged == true) .github/settings.yml: - New file for Probot Settings App - main branch: no required reviewers, no strict status checks, enforce_admins: false, allow force pushes, allow deletions https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
…dmin panel, disk protection - deploy.sh: fix early-return to check systemd service status before skipping deploy; if code is current but service is down, continue deploy to (re)start the service (PR 64 review fix) - index.js: username confirmation flow (WAITING_FOR_USERNAME never auto-accepts typed text; always shows "Yes, Continue / No, Edit" confirmation); NON_USERNAME_WORDS guard rejects obvious non-usernames (help, hi, back, ?, etc.) with a friendly nudge - index.js: AFFILIATE_REMINDER_TEXT shown at every required touchpoint: onboarding intro, link-username prompt, bonus request flow, promo claim view, weekly reminder, existing-account skip, and link-later flow - index.js: "Tell me more about the 30 SC bonus" button added to wagerReminderKeyboard, link-account prompt, bonus confirmation, and /start intro; w30_bonus_info action handler explains full promo with affiliate reminder and remaining attempts - index.js: admin panel expanded with "View Completed Requests", "Add Username Manually" (w30_admin_completed, w30_admin_link_username), "Return to Admin Menu" button, and inline "Mark Tip Sent" flow; finaliseUsernameLink helper centralises all username-save logic - index.js: submitBonusRequest includes attempt count, wager requirement, and affiliate reminder in user confirmation message; admin ping updated with "admin action required" label - scripts/disk-protect.sh: new weekly disk-space protection script (log rotation, compress logs >1 day, delete logs >7 days, journalctl vacuum 300 MB, npm cache clear, temp file cleanup, delete snapshots >30 days, never deletes .env or state files) - prod-run.sh: install disk-protect.sh as weekly Sunday 3 AM cron job https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
… command reference - COPY.ageGate: add "100% free to play, worldwide" statement with Markdown formatting; all senders updated to pass parse_mode: 'Markdown' - Intro GIF caption + post-age-gate rundown: add explicit "100% FREE / worldwide access" line at all onboarding touchpoints - userMainMenuText: add free-to-play reminder line in the main menu header shown to every user on every menu open - pmenu_help: changed from directly opening the booklet to a Help sub-menu with two options: "📖 Command Help Booklet" and "🐞 Report a Bug"; both sub-actions fully wired (help_open_booklet, help_open_bugreport) - adminMainMenuKeyboard: added "📖 Admin Commands Reference" and "🐞 Bug Reports" buttons directly on the admin main menu (not buried in a sub-menu); both wired to pamenu_admin_help and pamenu_bug_reports - configureBotSurface: complete rewrite of all three command scopes (global/default, all_private_chats, all_group_chats) plus the per-admin chat scope; every registered command now has an accurate description; added leaderboard_weekly, giveaway, and all admin commands (deploy, deploy_status, logs, admin_notify, etc.) - buildHelpPages page 1: add "100% FREE / worldwide" bullet, update quick-start to mention GambleCodez affiliate step and 30 SC bonus - buildHelpPages page 5: completely rewritten as a full user command reference grouped by category (Account, Runewager Account, Bonuses, Community, Help) with per-command tooltip text for every command - buildHelpPages page 6: completely rewritten as a full admin command reference grouped by category (Dashboard, 30 SC Bonus, Giveaway, Announcements, User Mgmt, Bug Reports, System) with usage notes - ageGateKeyboard: button labels updated to "Yes — I am 18+ and eligible" / "No — I am under 18" for clarity https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
deploy.sh — three fixes:
1. ERR trap recovery: trap now uses ${BASH_LINENO[0]} instead of
$LINENO (which resolves to the trap function's own line, not the
failing caller). A _SERVICE_STOPPED flag tracks whether systemctl
stop already ran; the trap attempts systemctl start on the existing
code so the bot is not left permanently down after a mid-deploy
failure. The flag is cleared whenever a successful start or rollback
start fires.
2. npm downtime risk: npm install / npm ci is now wrapped in the same
conditional guard pattern used for git fetch/reset. On npm failure
the script logs the full npm output, notifies admins, restarts the
service on the existing node_modules, and exits 1 — bot comes back
up rather than staying down.
3. Broad .env export: replaced `set -o allexport; source .env;
set +o allexport` with targeted grep extraction of only BOT_TOKEN
and ADMIN_IDS. All other .env variables are never exported into the
deploy environment, eliminating the risk of unintentional or
malicious variable leakage.
scripts/disk-protect.sh — two fixes:
4. Safety-check logic bug: CRITICAL_OK is now set to false (not left
as the initial true) when a critical file is found to be missing.
The subsequent `if [[ "$CRITICAL_OK" != "true" ]]` branch now
actually fires and emits the warning as intended.
5. Deletion logging suppression: all `find ... -delete -print | while
read` patterns replaced with a delete_found() helper that uses
`find ... -print0` + process substitution + `while IFS= read -r
-d ''` + `rm -f` in the loop body. This is reliable on both GNU
and BSD find, handles filenames with spaces/special chars, and
guarantees the log entry is written for every file actually deleted.
The same null-delimited pattern is applied to compress (step 2)
and /tmp cleanup (step 6).
https://claude.ai/code/session_01X3PxGFF5zzKptQwVkjYzzN
|
CodeAnt AI is reviewing your PR. |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Caution Review failedThe pull request is closed. 📝 WalkthroughWalkthroughTwo deployment and disk-protection shell scripts are enhanced with improved error handling, state tracking, and safety mechanisms. Deploy.sh adds pre-deployment checks, service recovery logic, and selective environment variable loading. Disk-protect.sh introduces a centralized file deletion helper function and critical file safety verification. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Poem
Note 🎁 Summarized by CodeRabbit FreeYour organization is on the Free plan. CodeRabbit will generate a high-level summary and a walkthrough for each pull request. For a comprehensive line-by-line review, please upgrade your subscription to CodeRabbit Pro by visiting https://app.coderabbit.ai/login. Comment |
Nitpicks 🔍
|
| journalctl --vacuum-size="${JOURNAL_VACUUM_MB}M" 2>&1 | grep -v "^$" | while read -r line; do | ||
| journalctl --vacuum-size="${JOURNAL_VACUUM_MB}M" 2>&1 | grep -v "^$" | while IFS= read -r line; do | ||
| log " [journal] $line" | ||
| done |
There was a problem hiding this comment.
Suggestion: Because the script uses set -euo pipefail, a non-zero exit status from journalctl --vacuum-size (for example, due to insufficient permissions) will cause the entire disk-protect script to abort at Step 4, skipping later cleanup and safety checks even though the journal vacuum is intended as a best-effort, non-critical operation; adding a guard to ignore failures here prevents unintended termination. [logic error]
Severity Level: Major ⚠️
- ⚠️ Disk cleanup script aborts at journal vacuum step.
- ⚠️ Later log compression and temp file cleanup skipped.
- ⚠️ Safety check for .env and runtime-state.json may skip.| done | |
| done || true |
Steps of Reproduction ✅
1. Install the weekly disk-protect cron job by running `/workspace/Runewager/prod-run.sh`,
which writes `0 3 * * 0 $DISK_PROTECT_SCRIPT >> ${PROJECT_DIR}/logs/disk-protect.log 2>&1`
into crontab (see `/workspace/Runewager/prod-run.sh`, disk-protect block around lines
22–32 in the shown snippet).
2. Ensure the cron (or manual) execution runs
`/workspace/Runewager/scripts/disk-protect.sh` as a non-root user that lacks permission to
vacuum the systemd journal (file `/workspace/Runewager/scripts/disk-protect.sh`, `set -euo
pipefail` at line 12 and Step 4 at lines 85–92).
3. When Step 4 executes, `journalctl --vacuum-size="${JOURNAL_VACUUM_MB}M"` (line 87)
exits with a non-zero status due to insufficient permissions; with `set -e` and `pipefail`
enabled, this non-zero pipeline exit causes the entire script to terminate immediately
because the pipeline is not guarded by `|| true` or an `if` condition.
4. Observe in `logs/disk-protect.log` that log messages for later steps (e.g., "Step 5:
Clear npm cache" at line 95, "Step 6: Clear temp files…" at line 104, safety check "Step
8" at line 124, and disk usage summary "Step 9" at line 139) are missing, confirming that
the script aborted at Step 4 instead of treating journal vacuum as best-effort.Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** scripts/disk-protect.sh
**Line:** 89:89
**Comment:**
*Logic Error: Because the script uses `set -euo pipefail`, a non-zero exit status from `journalctl --vacuum-size` (for example, due to insufficient permissions) will cause the entire disk-protect script to abort at Step 4, skipping later cleanup and safety checks even though the journal vacuum is intended as a best-effort, non-critical operation; adding a guard to ignore failures here prevents unintended termination.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.| if command -v npm >/dev/null 2>&1; then | ||
| npm cache clean --force 2>&1 | tail -1 | while read -r line; do log " [npm] $line"; done | ||
| npm cache clean --force 2>&1 | tail -1 | while IFS= read -r line; do log " [npm] $line"; done | ||
| log " npm cache cleared" |
There was a problem hiding this comment.
Suggestion: With set -euo pipefail enabled, a non-zero exit from npm cache clean --force will cause the whole script to exit at Step 5 even though clearing the npm cache is non-critical, and it will still log "npm cache cleared" on failure; wrapping this pipeline in an if and handling the failure explicitly avoids terminating the script and ensures success is logged only when the command actually succeeds. [logic error]
Severity Level: Major ⚠️
- ⚠️ Npm cache errors abort script before later cleanup steps.
- ⚠️ No dedicated error log for npm cache failures.| if command -v npm >/dev/null 2>&1; then | |
| npm cache clean --force 2>&1 | tail -1 | while read -r line; do log " [npm] $line"; done | |
| npm cache clean --force 2>&1 | tail -1 | while IFS= read -r line; do log " [npm] $line"; done | |
| log " npm cache cleared" | |
| if command -v npm >/devnull 2>&1; then | |
| if npm cache clean --force 2>&1 | tail -1 | while IFS= read -r line; do log " [npm] $line"; done; then | |
| log " npm cache cleared" | |
| else | |
| warn " npm cache clear failed (non-fatal)" | |
| fi |
Steps of Reproduction ✅
1. Ensure `/workspace/Runewager/scripts/disk-protect.sh` is executed with `set -euo
pipefail` enabled (line 12) either via the cron job installed in
`/workspace/Runewager/prod-run.sh` (disk-protect cron block around lines 22–32 in the
shown snippet) or by invoking the script manually.
2. Run the script on a system where `npm` is present (`command -v npm >/dev/null 2>&1` at
line 96 succeeds) but where `npm cache clean --force` will exit with a non-zero status
(for example, due to permission problems or a corrupted npm cache), causing the pipeline
at line 97 to fail.
3. When Step 5 executes, the pipeline `npm cache clean --force 2>&1 | tail -1 | while ...`
(line 97) returns a non-zero exit status; under `set -e` with `pipefail`, this unguarded
failing pipeline causes the entire `disk-protect.sh` script to terminate at Step 5.
4. Observe in `logs/disk-protect.log` that subsequent steps (Step 6 temp file cleanup at
lines 104–111, Step 7 backup pruning at lines 114–119, Step 8 safety check at lines
124–136, and Step 9 disk usage summary at lines 139–145) are missing, confirming that an
npm cache error aborted the disk-protect run instead of being treated as best-effort.Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** scripts/disk-protect.sh
**Line:** 96:98
**Comment:**
*Logic Error: With `set -euo pipefail` enabled, a non-zero exit from `npm cache clean --force` will cause the whole script to exit at Step 5 even though clearing the npm cache is non-critical, and it will still log "npm cache cleared" on failure; wrapping this pipeline in an `if` and handling the failure explicitly avoids terminating the script and ensures success is logged only when the command actually succeeds.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.|
CodeAnt AI finished reviewing your PR. |
CodeAnt-AI Description
Make disk cleanup reliable and add safer VPS deploy with recovery
What Changed
Impact
✅ Fewer accidental file-losses during cleanup✅ Shorter downtime after failed deploys✅ Clearer deploy failure and recovery notifications💡 Usage Guide
Checking Your Pull Request
Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.
Talking to CodeAnt AI
Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:
This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.
Example
Preserve Org Learnings with CodeAnt
You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:
This helps CodeAnt AI learn and adapt to your team's coding style and standards.
Example
Retrigger review
Ask CodeAnt AI to review the PR again, by typing:
Check Your Repository Health
To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.