Skip to content

Releases: iblai/infra-cli

v1.19.1

Choose a tag to compare

@bnsoni bnsoni released this 14 Aug 14:42
dbb4f23

[1.19.1] — 2026-08-14

Patch release. Recommended for anyone who has set or rotated an LLM API key with this CLI — see Upgrading.

Fixed

  • The LLM API key did not reach the platform. The credential row keeps the same secret in two columns — a plaintext one and an encrypted one — and current platform versions read the encrypted column as the source of truth, without falling back to plaintext. iblai infra llm set-key wrote only the plaintext column, so on a fresh environment the key stayed inert, and on a rotation the previous key remained in force — while the command reported success either way. The platform's own backfill does not repair a rotation: it fills the encrypted column only when that column is empty. Both columns are now written together with the same {"key": "..."} payload, and the columns are looked up on the model first so older deployments that never had the encrypted one still work.
  • Setting a key now clears is_preferred on the other providers. The platform takes the first preferred credential with no tie-break, so two preferred rows meant an arbitrary one won and a newly set key could be ignored.
  • An API key can no longer carry shell or Python syntax. The key is interpolated into a command that wraps a Python program, where a quote ends the Python literal, a double quote ends the shell string, and $(...) is substituted by the shell before the command runs. Keys are now checked against an allowlist of the characters providers actually issue — on the model, so the wizard, the flag and the .env path are covered at once — and asserted again in the role so the check does not depend on its caller. This matters most in CI, where the key comes from a secret store rather than from the operator running the command.
  • The key is no longer echoed if the task fails. It is interpolated into the command, which Ansible prints on failure, and into any CI log capturing that output.

Added

  • iblai infra llm set-key --provider openai|anthropic. The credential is named for its provider and the platform matches that name exactly, so the value is validated and lowercased rather than passed through. Interactive runs pick from a list; non-interactive runs default to openai, unchanged.

Upgrading

No migration, and existing environments are not repaired retroactively. Any environment whose LLM key was set or rotated through the CLI is still running on whatever is in its encrypted column — re-run set-key on this version to correct it:

uv tool upgrade iblai-infra     # or: uv tool install --force git+https://github.com/iblai/infra-cli
iblai --version                 # iblai v1.19.1
iblai infra llm set-key <name>

Test count: 929 passing.

Full changelog: v1.19.0...v1.19.1

v1.19.0

Choose a tag to compare

@bnsoni bnsoni released this 13 Aug 16:36
928226a

[1.19.0] — 2026-08-13

Adds iblai infra spa — run a customised copy of a deployed SPA alongside the original, on its own port and its own domain, without touching the one the platform depends on.

iblai infra spa clone <name>                 # prompts for source, name, domain
iblai infra spa list <name>                  # what's deployed, with ports
iblai infra spa remove <name> --spa <clone>

Also reachable from iblai infra configure <name>. Only the tagged Ansible role re-runs — no re-provisioning, no secret rotation.

Added

  • iblai infra spa clone <name> — picks the source from what is actually deployed on the server, allocates the next free port from 5060 (the stock SPAs hold 5000-5009), asks for the domain, and shows the whole plan before doing anything. --from, --as, --domain and --port skip the prompts.
  • iblai infra spa list <name> — what is deployed, with ports, marking which are stock and which are clones.
  • iblai infra spa remove <name> --spa <clone> — removes the containers, the nginx block and the directory. Refuses the platform's own SPAs, since the same role pointed at mentor or auth would delete the real one.
  • The clone copies the source's running environment file, not a re-render from config.yml, so it starts identical to what the source is actually serving including anything hand-edited on the box. PORT is then rewritten to the clone's own port — written rather than substituted, because deployments older than the template that introduced PORT have no line to replace, and a missing PORT leaves the clone listening on the source's port while compose publishes a different one: up, healthy-looking, and serving nothing. The clone is probed on its own port before the run is called a success.
  • Server blocks go in /etc/nginx/conf.d/custom_domains/, which the platform's proxy sync already excludes, so they survive ibl global-proxy regenerating everything else. The stock nginx.conf includes conf.d/*.conf without recursing, so the include for that subdirectory is added idempotently. nginx -t runs before every reload, since this happens against a live server.
  • Served over HTTP; put the domain behind whatever already terminates TLS for the environment.
  • Names and domains are checked against an allowlist before they are used. Both names become filesystem paths that the role acts on with elevated privileges, and the domain is written into an nginx server block, so a name has to be lowercase letters, numbers, hyphens and underscores, and a domain has to be a plain hostname. The same checks are asserted inside the roles, so they hold regardless of the caller.

Upgrading

No migration. Existing projects are unaffected — the two new roles are gated off and inert during a normal setup.

uv tool upgrade iblai-infra     # or: uv tool install --force git+https://github.com/iblai/infra-cli
iblai --version                 # iblai v1.19.0

Test count: 894 passing.

Full changelog: v1.18.1...v1.19.0

v1.18.1

Choose a tag to compare

@bnsoni bnsoni released this 10 Aug 17:23
4b31393

[1.18.1] — 2026-08-10

Patch release. Recommended for anyone on 1.18.0.

Fixed

  • Post-setup feature commands could report success without doing anything. On a call-server environment the tagged Ansible roles don't exist, and ansible-playbook exits 0 when a tag matches nothing — so iblai infra smtp enable would collect the mail settings, run zero tasks, and report "SMTP configured". These commands now refuse call-server deployments, and a run that matches no tasks is treated as a failure rather than success.
  • The SMTP port is validated at the prompt. A non-numeric port was previously only rejected after the rest of the form had been filled in, discarding every answer including the password.

Changed

  • Each post-setup feature has its own section in the README, covering what to prepare before running it — the OAuth redirect URI per SSO provider, Stripe live-mode behaviour, and which features restart services.

Test count: 843 passing.

v1.18.0

Choose a tag to compare

@bnsoni bnsoni released this 10 Aug 16:40
a471d72

[1.18.0] — 2026-08-05

Includes 1.17.0, which was never tagged.

Added

Optional integrations can be added after setup. SMTP, SSO, billing, an LLM key and extra tenants are all skippable during iblai infra setup — the credentials rarely exist on day one — but previously the only way to add one later was to re-run the whole playbook against a live environment.

iblai infra configure <name>          # menu of everything below

iblai infra smtp enable <name>        # outbound email (also: disable, status, enable-env)
iblai infra sso google <name>         # sign in with Google
iblai infra sso microsoft <name>      # sign in with Microsoft
iblai infra stripe enable <name>      # billing (also: enable-env)
iblai infra llm set-key <name>        # mentor service credential
iblai infra platform create <name>    # an additional tenant platform

Only the relevant Ansible role re-runs. None of them read the GitHub token or AWS keys, so adding a feature needs just the SSH key already recorded in the project plus that feature's own values.

Most take effect immediately. Google SSO, Stripe and the LLM key are database rows the platform reads per request. SMTP is the exception — it reaches the services as a container environment variable, so they must be recreated; the command says which and asks first (--no-restart, --yes). Microsoft SSO restarts Open edX because it changes settings read only at boot, and warns before doing so.

Changed

  • openai_api_key and admin_password are excluded from SetupConfig serialization, matching every other secret on the model.
  • The orphaned final_steps Ansible role is removed; its work had already moved to integrations, admin_setup and data_seeding.
  • CLAUDE.md brought current with the commands added since 1.14.0.

Test count: 839 passing.

v1.16.0

Choose a tag to compare

@bnsoni bnsoni released this 03 Aug 22:01
bea0095

[1.16.0] — 2026-08-03

Added

  • DNS verification. 1.15.0 started printing the records an operator has to create when the deployment does not manage DNS itself; it could not tell them whether those records were ever created correctly. Provisioning now offers to check, there is a re-runnable command, and setup warns before installing against domains that do not resolve.
    • Post-provision prompt — after apply on any externally-managed-DNS path, offers to verify the records now or skip and check later. Not shown when the stack created the records itself.
    • iblai infra dns check <name> — resolves every platform subdomain and reports, per record, whether it resolves and whether it points at this deployment's load balancer. A record that resolves somewhere else is reported as WRONG rather than passing, which is the case a plain reachability check misses. --watch re-checks on an interval until everything resolves, which suits waiting on a third party to publish the records. Exits non-zero while anything is unresolved, so it can gate a script.
    • Certificate state is reported alongside, since DNS resolving is only the first half: an ACM or Google-managed certificate cannot validate until the records exist, and PENDING_VALIDATION / PROVISIONING is what tells an operator they are waiting rather than broken. Best-effort and never fatal.
    • iblai infra setup <name> warns when the platform domains do not resolve, and asks before continuing. The platform routes by hostname, so installing early produces an environment that looks installed and then fails confusingly. --skip-dns-check bypasses it, and the check never blocks setup if it errors.
  • dnspython is now a dependency. Lookups go to public resolvers rather than the system one, so a stale local cache or split-horizon DNS cannot report success while the rest of the world still sees nothing — the whole point being to diagnose DNS an operator does not control.

Test count: 776 passing.

v1.15.0

Choose a tag to compare

@bnsoni bnsoni released this 23 Jul 19:02
b533d60

[1.15.0] — 2026-07-23

Added

  • Post-provision DNS instructions when no hosted zone is used. Terraform only creates DNS records automatically on the Route53 + ACM (AWS) and managed-existing-zone (GCP) paths. On every other path — an uploaded certificate, HTTP-only, or simply no matching hosted zone in the account — the load balancer is created but DNS is left to the operator, and the results screen previously showed only the load balancer address in a table row with no guidance. The provisioning summary now spells out exactly which records to create: for external DNS it lists a record for every platform subdomain pointing at the load balancer (CNAME → ALB DNS name on AWS, A → load-balancer IP on GCP), also written to dns-records.txt in the workspace, and notes that setup should run only once they resolve. When the stack created a new GCP Cloud DNS zone it now prints the zone's nameservers to delegate at the registrar (these were emitted as a Terraform output but never displayed). Route53/managed paths get a one-line confirmation that records were created. Applies to provision, provision-env, and launch (all share the results renderer).

Test count: 758 passing.

v1.14.0

Choose a tag to compare

@bnsoni bnsoni released this 23 Jul 17:37
0552812

[1.14.0] — 2026-07-23

Changed

  • Platform subdomain set updated and documented. Two backend data-API subdomains (mentor.data, web.data) were removed, and two application subdomains were renamed (mentorai → os, skillsai → lms). The change is applied consistently across every layer that references the subdomain set: DNS A-records and TLS certificate SANs (AWS single-server, AWS multi-server, and GCP templates), the IBL_SUBDOMAINS list in models.py, the application config that serves those endpoints, and the CSRF-exempt domain list. The full set is now listed in the README under "What gets created" and in docs/architecture.md.
  • Behavior change: environments created before this release will reconcile to the new DNS records and certificate on the next apply, and the two renamed application endpoints require a re-setup to take effect.

Test count: 750 passing.

v1.13.0

Choose a tag to compare

@bnsoni bnsoni released this 03 Jul 00:46
f5fb13a

Summary

Adds Google Cloud as a first-class provider for single-server deployments, alongside AWS — same wizard, same .env flows, same setup step.

  • Terraform stack (templates/gcp/single-server/): VPC + regional subnet, firewall rules (SSH restricted to the operator IP; health checks from Google's probe ranges), Compute Engine VM (Ubuntu 22.04, metadata SSH keys), unmanaged instance group behind a global external Application Load Balancer with a static IP, Google-managed SSL certificate (all platform subdomains, async validation), Cloud DNS A-records (existing zone auto-detected, or created with nameservers printed for registrar delegation)
  • Health check probes the LMS heartbeat with a learn.<domain> Host header — GCP accepts only a literal 200, and probing / hits the platform nginx catch-all's 301, marking the backend UNHEALTHY and serving 503 "no healthy upstream" for everything (hit live on the first bootstrap; designed out + regression-tested)
  • CloudProvider axis on InfraConfig (default aws; existing state files deserialize unchanged) + GCPCredentials (ADC or service-account key); runner dispatches templates/tfvars/env on it
  • PROVIDER=gcp non-interactive path + .env.provision.gcp.example; iblai infra permissions --provider gcp [--check]
  • One version prompt: setup/resetup ask for the prod-images release tag and resolve the matching iblai-cli-ops tag from its [tool.uv.sources] pin (uv ignores that table on git-URL installs, so the explicit install stays — only the question goes away). Stale 3.19.0 default removed from every input layer
  • Fixes: setup prompts crashed on GCP-provisioned states (no AWS credential block); iblai infra waf now cleanly rejects non-AWS stacks; GCP env builder surfaces validation errors instead of tracebacks
  • Docs: README rewrite (quick start, both clouds, ADC vs service-account walkthroughs, sample .env index) + step-by-step docs/GCP.md
  • Object storage remains on AWS S3 by design — operators supply credentials at the setup step (documented)

Test plan

  • 750 unit tests passing (GCP provider/runner/prompts/env suites mirror the AWS patterns)
  • terraform validate + fmt clean on the GCP templates
  • Verified live end-to-end on a real GCP project: provision (37 resources) → full 16-role bootstrap → platform serving over HTTPS behind the LB → teardown clean

🤖 Generated with Claude Code

v1.12.0

Choose a tag to compare

@bnsoni bnsoni released this 26 Jun 18:42
da4bb53

[1.12.0] — 2026-06-26

Fixed

  • TimescaleDB extension now created during DM bootstrap — the DM postgres image ships TimescaleDB preloaded (shared_preload_libraries=timescaledb), but the flow only ever ran CREATE EXTENSION vector; timescaledb was never created, so setup_timescale_views --full-setup (which ran under ignore_errors: true) silently degraded and analytics hypertables were never built. The ibl_dm role now runs CREATE EXTENSION IF NOT EXISTS timescaledb right after pgvector (idempotent, postgres-superuser). Confirmed against a field environment whose DB had only plpgsql + vector.
  • Microsoft SSO now uses the standard azuread-oauth2 backend — the microsoft_sso_config role derived the provider backend_name (and the /auth/login + /auth/complete SSO URLs) from platform_name (e.g. main-oauth2), which is not a registered Azure AD social-auth backend, so sign-in never completed and operators had to hand-fix the LMS OAuth2ProviderConfig. All backend references — OAuth2ProviderConfig.backend_name, the IBL_EDX.IBL_EDX_BASE_OAUTH_SSO_BACKEND block (IBL_OAUTH_SSO_NAME / TRACKED_PROVIDERS), other_settings.backend_uri, and IBL_SPA.AUTH.IBL_DIRECT_SSO_URL — now use the constant azuread-oauth2, matching the provider slug and the Azure-registered redirect URI. other_settings.platform_key still carries the tenant platform_name.
  • setup_timescale_views failures now surface — replaced the blanket ignore_errors: true on the data_seeding "Setup TimescaleDB views" task with a register + failed_when (fails on a genuine Python traceback, tolerates benign non-zero exits / idempotent re-runs) and a debug that prints the command output.

Added

  • IBL_DM.ENABLE_RBAC_GROUP_MANAGEMENT=true set by default — added to the ibl_platform "Enable DM RBAC" block alongside the existing ENABLE_RBAC / ENABLE_RBAC_SEEDING / ENABLE_TIMESCALEDB defaults (previously had to be set by hand).
  • Azure AD redirect-URI prerequisite documented — .env.setup.example and the microsoft_sso_config end-of-run confirmation now spell out the exact redirect URI the client must register in their Azure AD app: https://learn.<BASE_DOMAIN>/auth/complete/azuread-oauth2/.

v1.11.0

Choose a tag to compare

@bnsoni bnsoni released this 01 Jun 20:17
0fe62ac

[1.11.0] — 2026-06-01

Added

  • Optional AWS WAFv2 on the single-server ALB — opt-in at provision time via the wizard (a new sub-step of "Domain & Certificates" — default off), provision-env (ENABLE_WAF=true + WAF_ALLOWED_IPS=… in the .env), or launch / launch-env (--enable-waf + --waf-allowed-ips, or matching env keys). Attaches a Regional WAFv2 Web ACL to the ALB with rules tuned for ibl.ai's subdomain layout: admin-only allow rules (gated on an operator IP allowlist) for DM Swagger UI, edX Studio (CMS), Django /admin/, and DM /data; public allow rule for learn.<base> and apps.learn.<base>; six AWS managed rule groups (IpReputation, KnownBadInputs, Common, SQLi, WordPress, PHP); and a path-traversal block for .git / .env / .htaccess / .svn / .hg / .DS_Store. Total estimated WCU ≈ 1355 (under the 1500 default). Allowlist accepts both bare IPs (auto-suffixed /32) and CIDR.
  • iblai infra waf post-provision subgroup — toggle WAFv2 on an already-provisioned single-server stack without re-running the wizard. Four commands: enable [<name>] (interactive; on a project that already has WAF on, warns and prompts to update the allowlist with current IPs pre-filled), enable-env [<name>] -f .env (non-interactive, reads WAF_ALLOWED_IPS), disable <name> [--yes] (Y/N confirm by default; --yes for CI; removes the Web ACL + IPSet + association, leaves the ALB intact), status [<name>] (table of all WAF-eligible projects with no arg, detail panel with one). Rejects multi-server, call-server, bootstrap, and non-created projects up-front with a clear error. Subgroup module lives at src/iblai_infra/features/waf.py; the features/ package docstring documents the pattern for the next optional-feature toggles (SMTP, Stripe, SSO providers) so they can drop in with the same enable / enable-env / disable / status shape.
  • TerraformRunner.reapply() — shared helper for re-running Terraform on an existing workspace with the latest state.config. Re-copies .tf templates (so template fixes propagate), reads the existing terraform.tfvars to pin the original bucket_suffix (prevents accidental S3 bucket renames once the date-stamp window has rolled over), regenerates the rest of tfvars from state.config, then runs init → plan → apply. Returns parsed outputs. Used by both the new iblai infra waf <action> commands and the refactored iblai infra retry.
  • WAFv2 entries in REQUIRED_IAM_POLICY + a wafv2:ListWebACLs smoke check in check_permissions() so iblai infra permissions [--check] surfaces WAF readiness up-front instead of failing mid-apply.

Changed

  • _generate_tfvars(self, bucket_suffix: str | None = None) — accepts an optional pinned suffix. When None (today's default for first setup()), resolves the suffix from AWS as before. When provided, uses the pinned value. Load-bearing change for the new reapply() helper.
  • iblai infra retry now uses TerraformRunner.reapply() instead of inlining template-copy + init/plan/apply, removing a drift point with the new WAF subgroup. Behaviour is unchanged for operators: the existing failure-recovery guards and Route 53 CNAME conflict cleanup still run.