Releases: iblai/infra-cli
Release list
v1.19.1
[1.19.1] — 2026-08-14
Patch release. Recommended for anyone who has set or rotated an LLM API key with this CLI — see Upgrading.
Fixed
- The LLM API key did not reach the platform. The credential row keeps the same secret in two columns — a plaintext one and an encrypted one — and current platform versions read the encrypted column as the source of truth, without falling back to plaintext.
iblai infra llm set-keywrote only the plaintext column, so on a fresh environment the key stayed inert, and on a rotation the previous key remained in force — while the command reported success either way. The platform's own backfill does not repair a rotation: it fills the encrypted column only when that column is empty. Both columns are now written together with the same{"key": "..."}payload, and the columns are looked up on the model first so older deployments that never had the encrypted one still work. - Setting a key now clears
is_preferredon the other providers. The platform takes the first preferred credential with no tie-break, so two preferred rows meant an arbitrary one won and a newly set key could be ignored. - An API key can no longer carry shell or Python syntax. The key is interpolated into a command that wraps a Python program, where a quote ends the Python literal, a double quote ends the shell string, and
$(...)is substituted by the shell before the command runs. Keys are now checked against an allowlist of the characters providers actually issue — on the model, so the wizard, the flag and the.envpath are covered at once — and asserted again in the role so the check does not depend on its caller. This matters most in CI, where the key comes from a secret store rather than from the operator running the command. - The key is no longer echoed if the task fails. It is interpolated into the command, which Ansible prints on failure, and into any CI log capturing that output.
Added
iblai infra llm set-key --provider openai|anthropic. The credential is named for its provider and the platform matches that name exactly, so the value is validated and lowercased rather than passed through. Interactive runs pick from a list; non-interactive runs default toopenai, unchanged.
Upgrading
No migration, and existing environments are not repaired retroactively. Any environment whose LLM key was set or rotated through the CLI is still running on whatever is in its encrypted column — re-run set-key on this version to correct it:
uv tool upgrade iblai-infra # or: uv tool install --force git+https://github.com/iblai/infra-cli
iblai --version # iblai v1.19.1
iblai infra llm set-key <name>Test count: 929 passing.
Full changelog: v1.19.0...v1.19.1
v1.19.0
[1.19.0] — 2026-08-13
Adds iblai infra spa — run a customised copy of a deployed SPA alongside the original, on its own port and its own domain, without touching the one the platform depends on.
iblai infra spa clone <name> # prompts for source, name, domain
iblai infra spa list <name> # what's deployed, with ports
iblai infra spa remove <name> --spa <clone>Also reachable from iblai infra configure <name>. Only the tagged Ansible role re-runs — no re-provisioning, no secret rotation.
Added
iblai infra spa clone <name>— picks the source from what is actually deployed on the server, allocates the next free port from 5060 (the stock SPAs hold 5000-5009), asks for the domain, and shows the whole plan before doing anything.--from,--as,--domainand--portskip the prompts.iblai infra spa list <name>— what is deployed, with ports, marking which are stock and which are clones.iblai infra spa remove <name> --spa <clone>— removes the containers, the nginx block and the directory. Refuses the platform's own SPAs, since the same role pointed atmentororauthwould delete the real one.- The clone copies the source's running environment file, not a re-render from
config.yml, so it starts identical to what the source is actually serving including anything hand-edited on the box.PORTis then rewritten to the clone's own port — written rather than substituted, because deployments older than the template that introducedPORThave no line to replace, and a missingPORTleaves the clone listening on the source's port while compose publishes a different one: up, healthy-looking, and serving nothing. The clone is probed on its own port before the run is called a success. - Server blocks go in
/etc/nginx/conf.d/custom_domains/, which the platform's proxy sync already excludes, so they surviveibl global-proxyregenerating everything else. The stocknginx.confincludesconf.d/*.confwithout recursing, so the include for that subdirectory is added idempotently.nginx -truns before every reload, since this happens against a live server. - Served over HTTP; put the domain behind whatever already terminates TLS for the environment.
- Names and domains are checked against an allowlist before they are used. Both names become filesystem paths that the role acts on with elevated privileges, and the domain is written into an nginx server block, so a name has to be lowercase letters, numbers, hyphens and underscores, and a domain has to be a plain hostname. The same checks are asserted inside the roles, so they hold regardless of the caller.
Upgrading
No migration. Existing projects are unaffected — the two new roles are gated off and inert during a normal setup.
uv tool upgrade iblai-infra # or: uv tool install --force git+https://github.com/iblai/infra-cli
iblai --version # iblai v1.19.0Test count: 894 passing.
Full changelog: v1.18.1...v1.19.0
v1.18.1
[1.18.1] — 2026-08-10
Patch release. Recommended for anyone on 1.18.0.
Fixed
- Post-setup feature commands could report success without doing anything. On a call-server environment the tagged Ansible roles don't exist, and
ansible-playbookexits 0 when a tag matches nothing — soiblai infra smtp enablewould collect the mail settings, run zero tasks, and report "SMTP configured". These commands now refuse call-server deployments, and a run that matches no tasks is treated as a failure rather than success. - The SMTP port is validated at the prompt. A non-numeric port was previously only rejected after the rest of the form had been filled in, discarding every answer including the password.
Changed
- Each post-setup feature has its own section in the README, covering what to prepare before running it — the OAuth redirect URI per SSO provider, Stripe live-mode behaviour, and which features restart services.
Test count: 843 passing.
v1.18.0
[1.18.0] — 2026-08-05
Includes 1.17.0, which was never tagged.
Added
Optional integrations can be added after setup. SMTP, SSO, billing, an LLM key and extra tenants are all skippable during iblai infra setup — the credentials rarely exist on day one — but previously the only way to add one later was to re-run the whole playbook against a live environment.
iblai infra configure <name> # menu of everything below
iblai infra smtp enable <name> # outbound email (also: disable, status, enable-env)
iblai infra sso google <name> # sign in with Google
iblai infra sso microsoft <name> # sign in with Microsoft
iblai infra stripe enable <name> # billing (also: enable-env)
iblai infra llm set-key <name> # mentor service credential
iblai infra platform create <name> # an additional tenant platformOnly the relevant Ansible role re-runs. None of them read the GitHub token or AWS keys, so adding a feature needs just the SSH key already recorded in the project plus that feature's own values.
Most take effect immediately. Google SSO, Stripe and the LLM key are database rows the platform reads per request. SMTP is the exception — it reaches the services as a container environment variable, so they must be recreated; the command says which and asks first (--no-restart, --yes). Microsoft SSO restarts Open edX because it changes settings read only at boot, and warns before doing so.
Changed
openai_api_keyandadmin_passwordare excluded fromSetupConfigserialization, matching every other secret on the model.- The orphaned
final_stepsAnsible role is removed; its work had already moved tointegrations,admin_setupanddata_seeding. CLAUDE.mdbrought current with the commands added since 1.14.0.
Test count: 839 passing.
v1.16.0
[1.16.0] — 2026-08-03
Added
- DNS verification. 1.15.0 started printing the records an operator has to create when the deployment does not manage DNS itself; it could not tell them whether those records were ever created correctly. Provisioning now offers to check, there is a re-runnable command, and setup warns before installing against domains that do not resolve.
- Post-provision prompt — after
applyon any externally-managed-DNS path, offers to verify the records now or skip and check later. Not shown when the stack created the records itself. iblai infra dns check <name>— resolves every platform subdomain and reports, per record, whether it resolves and whether it points at this deployment's load balancer. A record that resolves somewhere else is reported asWRONGrather than passing, which is the case a plain reachability check misses.--watchre-checks on an interval until everything resolves, which suits waiting on a third party to publish the records. Exits non-zero while anything is unresolved, so it can gate a script.- Certificate state is reported alongside, since DNS resolving is only the first half: an ACM or Google-managed certificate cannot validate until the records exist, and
PENDING_VALIDATION/PROVISIONINGis what tells an operator they are waiting rather than broken. Best-effort and never fatal. iblai infra setup <name>warns when the platform domains do not resolve, and asks before continuing. The platform routes by hostname, so installing early produces an environment that looks installed and then fails confusingly.--skip-dns-checkbypasses it, and the check never blocks setup if it errors.
- Post-provision prompt — after
dnspythonis now a dependency. Lookups go to public resolvers rather than the system one, so a stale local cache or split-horizon DNS cannot report success while the rest of the world still sees nothing — the whole point being to diagnose DNS an operator does not control.
Test count: 776 passing.
v1.15.0
[1.15.0] — 2026-07-23
Added
- Post-provision DNS instructions when no hosted zone is used. Terraform only creates DNS records automatically on the Route53 + ACM (AWS) and managed-existing-zone (GCP) paths. On every other path — an uploaded certificate, HTTP-only, or simply no matching hosted zone in the account — the load balancer is created but DNS is left to the operator, and the results screen previously showed only the load balancer address in a table row with no guidance. The provisioning summary now spells out exactly which records to create: for external DNS it lists a record for every platform subdomain pointing at the load balancer (
CNAME→ ALB DNS name on AWS,A→ load-balancer IP on GCP), also written todns-records.txtin the workspace, and notes that setup should run only once they resolve. When the stack created a new GCP Cloud DNS zone it now prints the zone's nameservers to delegate at the registrar (these were emitted as a Terraform output but never displayed). Route53/managed paths get a one-line confirmation that records were created. Applies toprovision,provision-env, andlaunch(all share the results renderer).
Test count: 758 passing.
v1.14.0
[1.14.0] — 2026-07-23
Changed
- Platform subdomain set updated and documented. Two backend data-API subdomains (
mentor.data,web.data) were removed, and two application subdomains were renamed (mentorai→os,skillsai→lms). The change is applied consistently across every layer that references the subdomain set: DNS A-records and TLS certificate SANs (AWS single-server, AWS multi-server, and GCP templates), theIBL_SUBDOMAINSlist inmodels.py, the application config that serves those endpoints, and the CSRF-exempt domain list. The full set is now listed in the README under "What gets created" and indocs/architecture.md. - Behavior change: environments created before this release will reconcile to the new DNS records and certificate on the next
apply, and the two renamed application endpoints require a re-setup to take effect.
Test count: 750 passing.
v1.13.0
Summary
Adds Google Cloud as a first-class provider for single-server deployments, alongside AWS — same wizard, same .env flows, same setup step.
- Terraform stack (
templates/gcp/single-server/): VPC + regional subnet, firewall rules (SSH restricted to the operator IP; health checks from Google's probe ranges), Compute Engine VM (Ubuntu 22.04, metadata SSH keys), unmanaged instance group behind a global external Application Load Balancer with a static IP, Google-managed SSL certificate (all platform subdomains, async validation), Cloud DNS A-records (existing zone auto-detected, or created with nameservers printed for registrar delegation) - Health check probes the LMS heartbeat with a
learn.<domain>Host header — GCP accepts only a literal 200, and probing/hits the platform nginx catch-all's 301, marking the backend UNHEALTHY and serving 503 "no healthy upstream" for everything (hit live on the first bootstrap; designed out + regression-tested) CloudProvideraxis onInfraConfig(defaultaws; existing state files deserialize unchanged) +GCPCredentials(ADC or service-account key); runner dispatches templates/tfvars/env on itPROVIDER=gcpnon-interactive path +.env.provision.gcp.example;iblai infra permissions --provider gcp [--check]- One version prompt: setup/resetup ask for the prod-images release tag and resolve the matching
iblai-cli-opstag from its[tool.uv.sources]pin (uv ignores that table on git-URL installs, so the explicit install stays — only the question goes away). Stale3.19.0default removed from every input layer - Fixes: setup prompts crashed on GCP-provisioned states (no AWS credential block);
iblai infra wafnow cleanly rejects non-AWS stacks; GCP env builder surfaces validation errors instead of tracebacks - Docs: README rewrite (quick start, both clouds, ADC vs service-account walkthroughs, sample
.envindex) + step-by-stepdocs/GCP.md - Object storage remains on AWS S3 by design — operators supply credentials at the setup step (documented)
Test plan
- 750 unit tests passing (GCP provider/runner/prompts/env suites mirror the AWS patterns)
terraform validate+fmtclean on the GCP templates- Verified live end-to-end on a real GCP project: provision (37 resources) → full 16-role bootstrap → platform serving over HTTPS behind the LB → teardown clean
🤖 Generated with Claude Code
v1.12.0
[1.12.0] — 2026-06-26
Fixed
- TimescaleDB extension now created during DM bootstrap — the DM postgres image ships TimescaleDB preloaded (
shared_preload_libraries=timescaledb), but the flow only ever ranCREATE EXTENSION vector;timescaledbwas never created, sosetup_timescale_views --full-setup(which ran underignore_errors: true) silently degraded and analytics hypertables were never built. Theibl_dmrole now runsCREATE EXTENSION IF NOT EXISTS timescaledbright after pgvector (idempotent, postgres-superuser). Confirmed against a field environment whose DB had onlyplpgsql+vector. - Microsoft SSO now uses the standard
azuread-oauth2backend — themicrosoft_sso_configrole derived the providerbackend_name(and the/auth/login+/auth/completeSSO URLs) fromplatform_name(e.g.main-oauth2), which is not a registered Azure AD social-auth backend, so sign-in never completed and operators had to hand-fix the LMSOAuth2ProviderConfig. All backend references —OAuth2ProviderConfig.backend_name, theIBL_EDX.IBL_EDX_BASE_OAUTH_SSO_BACKENDblock (IBL_OAUTH_SSO_NAME/TRACKED_PROVIDERS),other_settings.backend_uri, andIBL_SPA.AUTH.IBL_DIRECT_SSO_URL— now use the constantazuread-oauth2, matching the provider slug and the Azure-registered redirect URI.other_settings.platform_keystill carries the tenantplatform_name. setup_timescale_viewsfailures now surface — replaced the blanketignore_errors: trueon the data_seeding "Setup TimescaleDB views" task with aregister+failed_when(fails on a genuine Python traceback, tolerates benign non-zero exits / idempotent re-runs) and a debug that prints the command output.
Added
IBL_DM.ENABLE_RBAC_GROUP_MANAGEMENT=trueset by default — added to theibl_platform"Enable DM RBAC" block alongside the existingENABLE_RBAC/ENABLE_RBAC_SEEDING/ENABLE_TIMESCALEDBdefaults (previously had to be set by hand).- Azure AD redirect-URI prerequisite documented —
.env.setup.exampleand themicrosoft_sso_configend-of-run confirmation now spell out the exact redirect URI the client must register in their Azure AD app:https://learn.<BASE_DOMAIN>/auth/complete/azuread-oauth2/.
v1.11.0
[1.11.0] — 2026-06-01
Added
- Optional AWS WAFv2 on the single-server ALB — opt-in at provision time via the wizard (a new sub-step of "Domain & Certificates" — default off),
provision-env(ENABLE_WAF=true+WAF_ALLOWED_IPS=…in the.env), orlaunch/launch-env(--enable-waf+--waf-allowed-ips, or matching env keys). Attaches a Regional WAFv2 Web ACL to the ALB with rules tuned for ibl.ai's subdomain layout: admin-only allow rules (gated on an operator IP allowlist) for DM Swagger UI, edX Studio (CMS), Django/admin/, and DM/data; public allow rule forlearn.<base>andapps.learn.<base>; six AWS managed rule groups (IpReputation, KnownBadInputs, Common, SQLi, WordPress, PHP); and a path-traversal block for.git/.env/.htaccess/.svn/.hg/.DS_Store. Total estimated WCU ≈ 1355 (under the 1500 default). Allowlist accepts both bare IPs (auto-suffixed/32) and CIDR. iblai infra wafpost-provision subgroup — toggle WAFv2 on an already-provisioned single-server stack without re-running the wizard. Four commands:enable [<name>](interactive; on a project that already has WAF on, warns and prompts to update the allowlist with current IPs pre-filled),enable-env [<name>] -f .env(non-interactive, readsWAF_ALLOWED_IPS),disable <name> [--yes](Y/N confirm by default;--yesfor CI; removes the Web ACL + IPSet + association, leaves the ALB intact),status [<name>](table of all WAF-eligible projects with no arg, detail panel with one). Rejects multi-server, call-server, bootstrap, and non-createdprojects up-front with a clear error. Subgroup module lives atsrc/iblai_infra/features/waf.py; thefeatures/package docstring documents the pattern for the next optional-feature toggles (SMTP, Stripe, SSO providers) so they can drop in with the sameenable / enable-env / disable / statusshape.TerraformRunner.reapply()— shared helper for re-running Terraform on an existing workspace with the lateststate.config. Re-copies.tftemplates (so template fixes propagate), reads the existingterraform.tfvarsto pin the originalbucket_suffix(prevents accidental S3 bucket renames once the date-stamp window has rolled over), regenerates the rest of tfvars fromstate.config, then runsinit→plan→apply. Returns parsed outputs. Used by both the newiblai infra waf <action>commands and the refactorediblai infra retry.- WAFv2 entries in
REQUIRED_IAM_POLICY+ awafv2:ListWebACLssmoke check incheck_permissions()soiblai infra permissions [--check]surfaces WAF readiness up-front instead of failing mid-apply.
Changed
_generate_tfvars(self, bucket_suffix: str | None = None)— accepts an optional pinned suffix. WhenNone(today's default for firstsetup()), resolves the suffix from AWS as before. When provided, uses the pinned value. Load-bearing change for the newreapply()helper.iblai infra retrynow usesTerraformRunner.reapply()instead of inlining template-copy + init/plan/apply, removing a drift point with the new WAF subgroup. Behaviour is unchanged for operators: the existing failure-recovery guards and Route 53 CNAME conflict cleanup still run.