-
Notifications
You must be signed in to change notification settings - Fork 32
console rps deplyment ADR
Proposal for deploying Console with existing RPS infrastructure (deployment#614).
- Abstract
- Console + RPS observations
- High-Level Design Decisions
- Detailed design by decision
- Appendix: problem backlog
A cloud deployment of Console and RPS runs both services side by side. RPS stays the provisioning orchestrator, and Console supports the RPS activation flow by serving device lifecycle APIs and managing devices once provisioned. This ADR proposes that RPS owns profiles, the two services keep separate databases and share one Vault, Console upserts devices and auth stays at the gateway.
The observations and what breaks during console and RPS services are run together.
Deployed the stack with make up, which brings up console, rps, db, vault, kong and webui. Local activation with Console works.
Remote mode, orchestrated by RPS. Every problem below is in this path.
| # | Problem | Effect | Decision |
|---|---|---|---|
| 1 | A profile created through Console's API is in the Sample Web UI, created through RPS's API is not listed | The UI is built "enterprise", which points to "rpsServer" at Console, so the profile list only ever reads Console's tables. Kong has no route to RPS's REST API, so "/rps" falls through to the UI's nginx. | 1 |
| 2 | Activation from the edge node fails with "ccm_profile_cira does not match list of available AMT profiles" | The profile row is there. RPS joins profiles_proxyconfigs on every profile read, and Console's migrations never create that table, so the query fails with "relation profiles_proxyconfigs does not exist". RPS swallows the error and returns no profile, so the device is told the profile does not exist. | 2 |
| 3 | Problem 2 fixed by creating the two missing tables at database init, activation fails with "Unknown error has occured" | RPS reads the AMT password from Vault at profiles/, and only RPS's own create writes it. The profile was created on Console, which encrypts the password into its own table, so the read 404s and RPS logs "password cannot be blank". The device is only told "Unknown error has occured". | 1 |
| 4 | Problem 3 worked around by seeding the Vault key, activation reports success but the device is not in the Sample Web UI | Console runs with AUTH_DISABLED=false and RPS sends no credentials, so all three POSTs to /api/v1/devices per activation — registration, CIRA registration and provisioning status — are 401 and no device row is ever created. Activation itself succeeds, and the only sign on the device is one line, "CIRA: Failed to save device to mps", followed by "Device activated successfully". | 4 |
| 5 | Problem 4 fixed by setting AUTH_DISABLED=true, the same three POSTs now return 400 and the device is still not registered | RPS puts the profile's tags straight into the device payload. Its own schema declares tags as text[], but Console's migrations created the shared column as text, so RPS reads back the string "rps-owned, cira" and Console's DTO expects an array. The body fails to bind before the insert is reached. Console logs the 400 with no reason. | 2 |
-
Remote mode provisioning stays in RPS, and RPS owns profiles. RPS keeps the activation, deactivation and maintenance state machines, and Console supports the activation flow by serving device lifecycle APIs and managing devices once provisioned.
RPS is the source of truth for profiles, CIRA, wireless, 802.1x, domain and proxy configs;
Console owns devices, AMT operations and login.
Console's own profile APIs are turned off by configuration. See 1.
-
Separate databases, one shared Vault. Console and RPS each keep their own database, because the two schemas are not compatible. They share one Vault, because RPS already writes the device secrets Console has to read. See 2.
-
Console must match the MPS device API, including upsert on POST. RPS registers the same device more than once per activation. Without upsert, CIRA provisioning fails. See 3.
-
The trust boundary is the network, not the device API. As in Cloud 2.x, auth is enforced at the gateway and the service APIs behind it are not, so Console runs with auth disabled and
8181reachable only from Kong and RPS. See 4.
RPS keeps activation, deactivation and maintenance. Console supports the activation flow by serving
device lifecycle APIs and managing devices once provisioned. No change to the RPS state machines. In
Docker, mps_server is set as RPS_MPS_SERVER=http://console:8181.
Ownership split
| Concern | Owner |
|---|---|
| Profiles, CIRA / wireless / 802.1x configs, domains, proxy configs | RPS |
| Provisioning orchestration, activation and deactivation | RPS |
| Devices, AMT operations, KVM/SOL, login | Console |
RPS is the only reader of a profile, at activation, and already holds the whole model including proxy
configs. Console-owned profiles would keep the two services joined at a schema, or need RPS to read
profiles over HTTP. One flow still crosses services: a CIRA config is created in RPS, but its
mpsRootCertificate comes from Console's /api/v1/ciracert.
Console still exposes profile CRUD on its own tables, so it must be turned off.
Proposal: add app.provisioning_owner, rps by default, console for Console standalone.
| Endpoint | console |
rps |
|---|---|---|
GET, POST, PATCH, DELETE on profiles, CIRA, wireless, 802.1x, domain configs |
as today | 403 |
| Devices, AMT operations, login | as today | as today |
GET /api/v1/server/features |
provisioningOwner: console |
provisioningOwner: rps |
RPS holds the full profile set once it owns provisioning, so Console's own copy is not needed for reads
either — GET is blocked along with the writes.
Console and RPS must not share a database. Both ship independent schemas for the same table names
(profiles, ciraconfigs, domains, wirelessconfigs, ieee8021xconfigs), with diverging column
types, and neither is a superset of the other. Console's DB_URL defaults to rpsdb today, so both
already migrate against the same database. RPS's init.sql no-ops against Console's CREATE TABLE IF NOT EXISTS, and a later Console migration fails on a duplicate column; Console's migrations never
create profiles_proxyconfigs, which RPS joins on every profile read.
Vault is already shared between MPS and RPS, and the namespaces don't collide. The only key Console
needs is devices/{guid}, which RPS writes at activation. MPSCerts has no producer once MPS is gone,
and RPS's TLS activation path reads MPSCerts.root_key.
Requirements:
| # | Requirement | Why |
|---|---|---|
| 1 |
rpsdb for RPS, consoledb for Console |
the schemas conflict either way round |
| 2 | Console's Compose file stops defaulting to rpsdb
|
that is the RPS name |
| 3 |
consoledb created at initdb |
the stack's Postgres creates rpsdb only |
| 4 |
SECRETS_PATH=secret/data on Console, no trailing slash |
Console joins base path and key with /; a trailing slash makes every read miss |
| 5 | One-shot seed of MPSCerts from Console's certs/root
|
RPS's TLS activation path reads it and nothing writes it now |
| 6 | Console's Vault token needs read on devices/{guid} only |
it shares RPS's token today, which is the root token |
The stack's init container only generates TLS certs and templates Kong config; #5 needs new logic
there, not just a config change.
Console must serve the endpoints RPS calls, and persist status, activatedAt and
lastProvisionedAt in deviceInfo:
GET /api/v1/devices/{guid}?tenantId=...POST /api/v1/devicesPATCH /api/v1/devicesDELETE /api/v1/devices/{guid}?tenantId=...DELETE /api/v1/devices/refresh/{guid}
The contract change is upsert: a POST for a GUID that already exists must update that row instead of
failing. MPS behaves this way; Console returns 409. RPS posts the same device more than once per
activation — registration and provisioning status always, plus CIRA registration with mpsusername on
the CIRA path — so a 409 on any of these blocks provisioning: the state machine goes to FAILURE,
mpsusername is never set, and Console never knows to use the tunnel.
Proposal: on GUID conflict, POST merges the supplied fields the same way PATCH does. 201 stays
"created", 200 means "merged". A caller could now overwrite mpsusername, hostname or tags on an
existing device; the real control against that is the trust boundary in
4.
DELETE /api/v1/devices/refresh/{guid} is called after an AMT password change. Console caches a WSMAN
connection per GUID; without this route the cached connection keeps using the old password. Adding the
route drops the cached connection so the next call rebuilds it with the password now in Vault.
The proposal keeps the Cloud 2.x trust model: auth is enforced at the gateway, not on the service APIs
behind it. MPS mounts /api/v1 with no JWT middleware and RPS→MPS calls carry no credentials, so
Console runs AUTH_DISABLED=true here too. Login still works, since /api/v1/authorize sits outside
the JWT-protected group.
This rests on one invariant: 8181 is only on the internal Compose network, never published to the
host. If it were published to an untrusted network, an unauthenticated caller would get the full
device API — merge, delete and AMT operations.
Known gaps, both inherited from Cloud 2.x and bounded by that invariant:
- RPS→Console is unauthenticated. Any caller on the internal network can register, merge or delete a device, or drive AMT operations, indistinguishably from RPS. Closing this needs a service token or mTLS between RPS and Console — an RPS change, out of scope for v3.0.
-
The relay loses Console's own token check. MPS verifies the JWT on every relay upgrade
independently of auth mode; Console's check is wrapped in
if !auth.disabled, so with auth disabled KVM/SOL relay is guarded by Kong alone. Proposal: make Console's relay check independent ofauth.disabled, since that token is device-scoped, not an API session.
Not yet reproduced. Moved into Console + RPS problems one at a time as each is confirmed.
| Problem | Effect | Decision |
|---|---|---|
Both services own profiles, and the UI's enterprise build sends profile writes to Console |
RPS activates against its own tables, so a profile created in the UI is not the one used | 1 |
Console and RPS share rpsdb, with conflicting schemas for the same table names |
whichever service initialises the database breaks the other; profiles_proxyconfigs is never created and RPS joins it on every profile read |
2 |
SECRETS_PATH=secret/data/ has a trailing slash |
Console joins base path and key with /, so every Vault read misses |
2 |
MPSCerts has no producer once MPS is gone |
RPS's TLS activation path reads MPSCerts.root_key
|
2 |
POST /api/v1/devices returns 409 for a GUID that already exists |
RPS posts the same device more than once per activation, so CIRA provisioning fails | 3 |
Console has no DELETE /api/v1/devices/refresh/{guid}
|
after an AMT password change the cached WSMAN connection keeps the old password | 3 |
Console runs with AUTH_DISABLED=false behind the gateway |
RPS sends no credentials, so every device registration is 401 and no device is created | 4 |
| Activation: profile not found |
proxyconfigs and profiles_proxyconfigs are in RPS's init.sql but nothing ever applies it, so RPS's profile query dies on a missing relation |
2 |
| Activation: password cannot be blank | RPS reads the AMT password only from Vault, and only RPS's own create writes it, so Console-created profiles can't activate | 1 |
| No route to RPS's REST API |
/rps matched nothing and fell through to the UI's nginx, which answers GET with index.html 200, producing a phantom "profile already exists" |
1 |
| The two creation paths are mutually exclusive | Console-created profiles show in the UI but can't activate; RPS-created ones activate but break the UI list | 1 |
| RPS can't authenticate to Console |
saveDeviceToMPS posts to /api/v1/devices with no Authorization header, so CIRA registration 401s |
4 |
AUTH_DISABLED can't satisfy both |
true breaks UI login, false breaks RPS's device save, and no deployment-config setting fixes it |
4 |