Repository navigation
fix(security): resolve CodeQL alerts blocking the v3 release - #300
Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…on, restrict CI token
* feat: Gordon v2.50.0 application deployment model (#276) * feat(apps): add declarative application deployments * fix(security): enforce host trust boundaries * fix(security): close remaining host trust gaps * test(security): align hardened image fixtures * refactor(cli): route local commands through admin socket * fix(admin): authorize local control plane routes * feat(install): build next channel from pinned source * feat(install): default to user-local binary path * fix(ci): avoid overlapping Go module caches * fix(apps): address wrap-up review findings * fix(apps): close v2.50 deferred findings (#280) * fix(apps): persist mutation identity and close state cleanly * refactor(cli): unify app control plane capabilities * fix(apps): claim mutation keys atomically Add a durable request identity to app operation journals and one atomic store claim that binds an idempotency key to exactly one request: - domain.AppOperation carries the immutable request identity (kind, app, requested revision, requested service). - out.AppState.ClaimOperation persists the in-flight journal for an absent key, replays the stored journal for the same request, and refuses a key that already answered a different request with ErrAppStateConflict. - A claim for a name with no live app identity fails with the new ErrAppNotFound and writes nothing; the admin API maps it to 404 app-not-found. read-only AppExists backs the not-found normalization for revisions. - Every keyed deploy/stop/start/restart/remove claims before any effect. A replayed key returns its stored journal and runs no workload call; an unfinished journal returns the stored record with a conflict. - The apps use case returns the journaled operation together with the engine error instead of dropping the record. * fix(apps): order operation steps behind publication and complete read models Journal ordering: - A deploy service step is recorded as succeeded only after the runtime effect, the ACTIVE publication, and the traffic graph apply succeed; a rejected apply fails the step instead of leaving a success behind. - Lifecycle verbs record a failed traffic.publish step and a failed outcome when the final graph apply is rejected, so a rejected publication never replays as success. - Every error path after the claim now records a terminal failure, so a keyed repeat replays the failure instead of reading a stuck claim. - start marks a service step succeeded only after its effective state is published. Read models: - Show returns not-found for a name with no live app identity, and reports desired acceptance/pending status, per-service effective revision, container, digest, restart safety, owned resources (volumes, secret paths, image references), and the latest operation with kind, outcome, and start time. - List skips retired names and reports desired status, pending, convergence, stop intent, and the last outcome. Pending is derived from desired vs ACTIVE, never from a journal. - Admin DTO mappings carry these fields; deploy responses attach current effective and owned state instead of inferring convergence from the journal outcome. The unused observed field is gone. * fix(app): publish the host index only after the graph apply succeeds Traffic publication now prepares and commits instead of mutating the live host index first: - ACTIVE state projects into a detached candidate index that the graph is built and applied against. The live index is swapped only after the dataplane accepted the graph, so a rejected apply leaves both the index and the graph at the previous generation. - A failing projection applies nothing at all. - Withdrawal stays fail-closed: when the graph apply fails after the durable binds were cleared, the read model still drops the withdrawn host so the proxy never dials a recycled loopback port; the error is returned unchanged so the caller keeps its publication inhibition. HostIndex gains an Entries accessor for the swap, and the publisher takes its graph applier as a wired dependency so the ordering is testable without a live traffic manager. * fix(runtime): retire exact containers with the effective stop grace Runtime contract: - StopContainer and RestartContainer take the effective stop grace. The Docker adapter converts it to whole API seconds, rounding up, and keeps the runtime default for a non-positive grace instead of an immediate kill. The fixed twenty-second timeout is gone. - Every app path passes the grace of the generation it stops: stop, restart, removal, replacement retirement, readiness-failure cleanup, and recovery cleanup. A missing declared grace falls back to the app default, never to zero. One retirement path: - retireContainer stops with grace (or force-removes a never-published candidate), confirms the exact ID is gone, and only then releases the container's backend claims and clears its recovery inhibition. A container that is not confirmed gone keeps its claims, its inhibition, and its ACTIVE reference, and fails the operation. - Stop, remove, deploy service removal, HTTP replacement retirement, interrupted replacement, readiness-failure candidate cleanup, create and start cleanup, backend-bind registration failure, and container-not-found recovery all use it, so no disappearance path leaks a loopback claim. - A confirmed-missing container releases its claims during recovery. Bounded leftovers: - Journaled operations carry bounded cleanup warnings (service, leftover, detail) for anything that could not be removed or released; they reach the admin response instead of being dropped, and never contain logs. * fix(apps): reclaim private app networks on removal Removal now removes the app's incarnation-owned private networks immediately, once every exact container is confirmed gone and while the ownership record is still live: - The runtime name is re-derived from the incarnation UUID and the observed labels must prove that exact ownership, so a name reused by a later incarnation can never delete the previous one's network. - A network with attached containers, or with labels that do not prove this incarnation, is left in place and reported as a bounded warning. - A failed removal aborts the operation before the incarnation is retired, so the stopped intent, the inhibition, and ACTIVE state stay and a retry can converge. - Shared networks are never removed by an app removal, and retained volumes, secrets, and images are untouched. * refactor(backup): resolve app backups from manifest declarations Backup targets are now read from the app's ACTIVE record instead of being inferred, and the app name is the backup identity: - A target exists because a service declares it: a PostgreSQL database in service.databases referenced by service.backup.postgres, or a volume in service.volumes referenced by service.backup.volume. The image, port, container label, and attachment heuristics are gone, along with the detection endpoint and DTOs that exposed them. - Selectors are explicit: an omitted --service or resource selector only succeeds when exactly one compatible target exists, ambiguity lists the candidate service/resource names, and an unknown target is not found. - The scheduler runs only databases whose own declared schedule matches the firing tier, and volume runs follow the installation preset. - Artifacts are stored under the app name: filesystem paths and S3 keys use apps/<app>/..., never a domain prefix. Existing domain-keyed artifacts are left untouched and are not read. - The admin API and CLI rename domain fields to app, add service, and name the resource database or volume. The CLI gains `backup run APP --service S --database D` and `backup volume run APP --service S --volume V`; the domain-keyed `detect` command and the databases/volumes aliases are gone. * refactor(cli): unify the app and daemon control planes The CLI now has one ControlPlane interface and one daemon-backed implementation: - App methods moved into ControlPlane; AppControlPlane, remoteAppControlPlane, and app_controlplane.go are gone, and remoteControlPlane delegates every operation to a single remote.Client (explicit remote or owner-only local admin socket). - App commands resolve through resolveAppPlane, which returns a handle they close, instead of a separate resolver with no ownership. - CLI tests drive the generated MockControlPlane: the handwritten fakeAppPlane and status shim are gone, and the source-text parity and UI adoption checks are replaced by the behavior they guarded (local socket and explicit-remote resolution, DTO parity across transports, presentation helpers). * refactor: drop dead app plumbing and test-only production hooks - AppServiceImpl keeps one configured core apps Service instead of rebuilding it on every apply; the listener, barrier, and image-policy options configure that instance in place. - Remove the container cache/drain/metrics hooks that only stored their dependencies, the proxy container-deployed invalidation handler and its event payload, the container event constants nothing published or read, the internal-deploy context flag with no reader, the label constants the declarative cutover stopped using, and the deprecated revision retention alias. - Move the app-state corruption hooks out of the production store into export_test.go. * docs: describe current backup, log, and app identity behavior - Backup docs document declarative targets (app, service, database or volume), the explicit selector rules, app-keyed artifact paths, and the app-based JSON fields. - App docs describe the idempotency contract, the complete show read model, and app/service-only log identity. - `gordon logs` reads daemon process logs only: the domain-keyed container-log path and its control-plane methods are gone, and app workload output is read through `gordon apps logs`. * fix(apps): close deferred review gaps * feat(apps): fix migration friction and add internal HTTP interfaces (#282) * fix(apps): reduce migration friction * feat(apps): add internal HTTP interfaces * fix(apps): harden internal deployment lifecycle * fix(apps): preserve targeted deployment state (#283) * fix(apps): preserve networks in active diffs * fix(apps): allow safe targeted deployments * fix(apps): preserve non-targeted services * fix(apps): parse shared network tables * refactor(deployment): use durable sequential app deploys (#284) * refactor(deployment): use sequential service replacement Replace the rolling-versus-interrupted split for app services with one sequential replacement applied to every service: preflight -> withdraw traffic (confirmed) -> retire the superseded container (confirmed gone) -> create and start the replacement -> readiness probe -> publish ACTIVE -> rebuild traffic A deploy may briefly interrupt a service, and two Gordon-managed generations of one service never run at the same time. * refactor(deployment): drop httpEligible, deployHTTP, deployInterrupted, cutoverService, retireContainerID, drainWithDeadline and the retire-after-publish path * refactor(deployment): restart always restarts the same pinned container in place (withdraw, restart, readiness, republish) and no longer creates a second container * refactor(deployment): resolve the implicit HTTP readiness probe from the service spec so readiness no longer depends on how a service is replaced, keeping the check for stateless public HTTP services * refactor(deployment): remove DeployResult.Interrupted and ServiceResult.Retire/RetireGrace * test(deployment): replace the rolling-behaviour tests with sequential replacement, in-place restart, readiness-ordering, withdrawal-failure, cancellation, recovery-inhibition and multi-service failure tests * docs: describe sequential replacement, the expected short interruption, preflight protection and manifest-based rollback * fix(deployment): recover interrupted replacements safely Close the crash window left by sequential replacement and the recovery dead ends the review of it exposed. * fix(deployment): journal the replacement's container ID before it is started, and keep it in the journal when a replacement fails after creating its container, so a leftover stays traceable * fix(deployment): reconcile an unfinished operation at boot and before every mutation of an app: remove a never-published candidate (releasing its loopback claims) before creating or rebuilding a generation, keep the published generation, retry the leftovers of an operation that already failed, and finalize the operation (failure, or success when every step had already completed) so a keyed retry never reads a permanent in-flight claim * fix(deployment): refuse a new journal while a different operation of the same app is unfinished, so a claim can never mask an interrupted predecessor * fix(deployment): a restart that rebuilds a missing generation clears the recovery inhibition it wrote for the container proven gone, which otherwise refused every later start, recovery pass, and restart * refactor(deployment): recover the store before reconciling in deploy, move the preflight store recovery to its callers, and extract runtimeImageRef, convergeCandidate, and interruptedOutcome * refactor(config): remove the dead [deploy] configuration surface (pull_policy, readiness_*, health_timeout, stabilization_delay, probe timeouts, drain_*) from the viper defaults, the example configuration, and the docs: no code ever read these keys * test(deployment): cover the candidate journal write before start, boot and mutation-time reconciliation, terminal-operation leftovers, an interruption after the last step, the claim guard, restart rebuilding a missing container, and the inhibition cleanup * docs: document crash recovery, the 15 s republish cadence, and restart rebuilding from the pinned digest * fix(deployment): fail closed on orphan replacements and repair recovery Close the recovery findings left by the sequential-replacement review. * fix(deployment): fail closed when a replacement candidate cannot be removed. Reconciliation returns an error, keeps the journal (and its latest position) instead of finalizing it, and no later mutation proceeds, so a newer operation can never mask an orphan that may still run. * fix(deployment): clear the superseded container's stale replacement-pending inhibition once its never-published candidate is confirmed gone, so boot/start/restart can rebuild the generation ACTIVE records. The inhibition is kept while candidate cleanup fails, preserving the single writer. * fix(deployment): set Service on start/restart placeholder journal steps so a recovered step is attributable and its inhibition can be cleared. * fix(deployment): repair the Preflight(key)->Deploy(same key) contract: Preflight reconciles before resolving and never claims the caller's request key, so a later Deploy of the same key executes instead of replaying the preflight journal as a conflict. * refactor(deployment): drop the unused bool result of verifyRunningService. * test(deployment): cover fail-closed candidate removal, stale inhibition removal with a volume-owning rebuild, the repaired preflight/deploy key contract, and preflight reconciliation of an unfinished predecessor. * docs: describe fail-closed reconciliation and inhibition removal. * refactor(deployment): split deploy engine into start and execute phases Factor the synchronous deploy engine into a durable two-phase internal API. StartDeploy runs under the app lock: it reconciles any interrupted predecessor, resolves the revision and runs the convergence checks, atomically claims the request key, and persists a non-terminal journal. It performs no image pull, pinning, or workload mutation, and reports whether the caller newly owns the operation or is an idempotent replay. ExecuteDeploy reacquires the app lock, loads and verifies the claimed journal by app/op identity, runs image pull/pinning and the existing sequential replacement, and records the terminal outcome. A terminal or mismatched claim is never executed: it replays instead. Cancellation or shutdown mid-execution leaves the non-terminal journal for reconciliation. Service.Deploy stays the synchronous composition of both phases, so all existing callers and behavior are unchanged. Existing preflight is re-expressed as claimDeploymentLocked plus pinPreflightLocked. * feat(apps): orchestrate two-phase deploy in the background AppServiceImpl.Deploy now claims and persists the operation through StartDeploy and, for the key owner only, schedules ExecuteDeploy on a daemon-owned context so the request returns the running journal promptly and request cancellation never aborts the replacement. Duplicate keys replay their stored journal and never spawn a second execution. A daemon context plus WaitGroup on AppServiceImpl gives the lifecycle its graceful shutdown hook: Shutdown cancels in-flight executions, which leave the persisted non-terminal journal for reconciliation, and waits for them to unwind. * feat(app): wire daemon lifecycle context into app administration Build AppServiceImpl through newAppDaemonService with the daemon/supervisor lifecycle context, so background deploy executions are owned by the daemon and never by a request context. Graceful shutdown registers AppServiceImpl.Shutdown on the bounded 30s shutdown context as Phase 0.5, cancelling and joining in-flight deploy executions before traffic, runtime, or the app state store is torn down. A failed unwind is logged and never aborts teardown. * feat(apps): answer 202 for running deploys and quiesce kernel app admin The admin deploy handler now answers 202 Accepted with the durable running journal when it newly owns the operation, so the request returns promptly and the client observes completion via GET operations/by-key. A terminal replay stays 200, and conflicts, preflight failures, and failed replays keep their existing 409 mapping. AppDeployResponse carries a stable status field (running until terminal, then the persisted outcome); the by-key lookup keeps answering 200 and remains the recovery source. Permission checks and Idempotency-Key validation are unchanged, and the remote client's generic 2xx path already accepts 202. Kernel.Close now cancels and joins the daemon-owned AppServiceImpl on a bounded context before any remaining kernel resource is torn down, so a background deploy execution cannot outlive the state and runtime it uses. A nil implementation is never stored in the lifecycle interface. * feat(apps): poll 202 deploy operations and add operations watch `gordon apps deploy` now follows a backend 202 through the existing operations/by-key endpoint until the operation is terminal, printing concise progress only when operation or step state changes. JSON mode keeps stdout to a single final document (progress goes to stderr) and terminal partial/failed outcomes exit nonzero. Add `gordon apps operations watch APP --key KEY` to resume polling an existing operation without reissuing the mutation. Ctrl-C stops only local polling, never the daemon-side operation, and prints the resumable command in human mode. 404 fails fast; transient polling errors are bounded; local and remote planes share the same client path. Docs: apps deploy/operations sections and CLI index. * fix(deployment): keep live deploy claims out of foreground reconciliation Track the operations this process owns between the StartDeploy claim and the ExecuteDeploy exit, so a foreground mutation landing in the hand-off window cannot reconcile and terminalize a live claim. ExecuteDeploy now always reloads the operation from the store by app/op identity and refuses a terminal or mismatched record, never executing a stale in-process journal. * fix(apps): settle deploy claims raced by shutdown before scheduling Gate AppServiceImpl.Deploy on the daemon lifecycle before StartDeploy, so a request arriving after shutdown began is refused with a wrapped domain.ErrAppStateConflict and never claims a journal or answers 202. Enroll the Deploy hand-off with the lifecycle so Shutdown waits for an in-flight claim to be either scheduled or settled, closing the window where shutdown starts after StartDeploy but before scheduleExecution. When shutdown wins that race, settle the owned claim instead of dropping it: add deployment.Service.AbandonDeploy, which releases the in-process live marker and converges the still non-terminal journal as interrupted, so no accepted durable claim stays marked live forever and a later claim recovers cleanly. Add deterministic tests for stopping before a deploy and for shutdown racing between the claim and scheduling, plus an engine test proving the settled claim is terminal and never blocks the next owner. * fix(apps): reuse deploy watch for accepted apply --deploy `apps apply --deploy` now follows a 202 accepted deploy through the same by-key watch as `apps deploy`: progress only when the operation or step state changes, one terminal journal, and a nonzero exit for a terminal partial/failed outcome without hiding the successful apply. JSON mode emits exactly one final combined {apply,deploy} document and sends progress, transient warnings, and Ctrl-C resume guidance to stderr; no initial running document is written and the deploy is never re-POSTed. The command now threads cmd.ErrOrStderr so stderr routing is exercised. Add focused tests for 202 to success, partial/failure with preserved apply, unchanged-state deduplication, JSON purity, single mutation, and Ctrl-C resume in both human and JSON modes. * fix(apps): return journal shape for running deploy replays A same-key deploy replay whose operation is still in flight now always answers 409 with the stored operation journal, before or after any service step has started. Only an absent or terminal journal may map a preflight error envelope, so a non-terminal replay never degrades to the mapped error response. * refactor(deployment): remove dead exported Service.Preflight Remove the production-dead deployment.Service.Preflight entry point and its now-unreachable preflightLocked wrapper. No production caller and no boundary interface exposed it; only direct tests did. The private preflight gates (claimDeploymentLocked/pinPreflightLocked) used by StartDeploy/ExecuteDeploy are preserved. Delete the tests and helpers that existed only for the removed entry point. * test(deployment): restore preflight gate coverage through deploy phases Retarget the live preflight gates that were only reachable through the removed Service.Preflight onto StartDeploy/ExecuteDeploy and the synchronous Deploy composition: missing required secret (ErrAppSecretMissing), env/secret collision (ErrInvalidAppSpec), unmanaged image volume (ErrAppUnmanagedImageVolume), R2 unowned existing volume refusal (ErrAppStateConflict), the targeted-deploy diff policy, and the shared GC barrier held across resource selection. Every failing gate proves no workload mutation before the failure, and the execute-phase gates prove a terminal failed journal. No production code changes. * fix(deployment): address HTTP readiness at declared container port qBittorrent rejects the readiness probe with 401 "Invalid Host header, port mismatch" because the probe dials the random loopback published port and Go derives the Host header from it. Keep dialing the published bind but send 127.0.0.1:<declared container readiness port> as the Host authority, matching what the workload listens on. Thread the Host authority through the injectable httpGet dependency so tests can assert both the dial URL and the Host header, and cover the mismatch with a server that rejects published-port authorities. * fix(deployment): harden sequential deploy lifecycle Address the review findings on the sequential service replacement work: - keep the revision resolved at claim time authoritative for the whole execution instead of re-resolving desired state in the second phase - record successful candidate convergence by clearing the step After and keeping the container id in Detail, so a converged step is not retried while its audit trail survives - normalize the persisted service-step plan when a claim is resumed and converge recorded candidates before the plan is replaced - fail closed when app administration does not quiesce: skip state and runtime teardown and propagate the error to the process entry point - carry app and op fields through zerowrap contexts for the deploy scheduling logs instead of the service logger - move the operations watch command into apps_operations_watch.go - exercise the new admin deploy handler cases over a real httptest server * feat(apps): add allowlisted CDI devices for app services (#285) * feat(apps): add allowlisted CDI devices for app services Manifests declare logical devices via devices = [...]; the installation maps each name to explicit CDI IDs plus exact app/service allowlists under [app_devices.<name>]. Resolution happens at activation time and encodes one native CDI DeviceRequest; revisions persist logical names only. Revoked grants and unsupported engines fail closed before any workload mutation. * refactor(apps): extract helpers to satisfy gocyclo budget Split device diff entry, preflight error tail, and reload policy validation out of diffService, isMappedPreflightError, and applyLoadedConfig. * fix(apps): address device review findings Probe engine CDI capability in preflight via a new ContainerRuntime.SupportsCDIDevices boundary method so unsupported engines fail before workload mutation. Preserve ErrRuntimeUnsupported through createContainer redaction and map it to the structured runtime-unsupported envelope. Stop echoing resolved CDI IDs in duplicate-grant errors and run union resolution at apply time. * fix(apps): harden device error sanitization Hoist the runtime-unsupported sentinel above the bind redaction branch so mixed bind+device creates keep the structured envelope. Sanitize the preflight engine probe cause so daemon endpoints never reach journals or API responses. * fix(apps): harden device policy, engine gate, and probe errors Address review findings on the CDI device work: - docs: move `devices` above the first nested service table so TOML assigns it to the service, not to the last bind entry. - domain: keep the ErrInvalidAppSpec chain through device policy allowlist validation, and reject CDI device names that are host paths or raw device nodes instead of valid CDI names. - docker: parse engine major/minor strictly, so prerelease suffixes such as "-rc1" fail closed instead of clearing the version gate; keep cancellation and deadlines from the /version probe while dropping the cause text that embeds the daemon endpoint. - deployment: preserve cancellation and deadlines from the capability probe and collapse every other cause to ErrRuntimeUnsupported alone. - appstate: load the legacy no-devices record from a verbatim fixture instead of writing through the current serializer. * fix(traffic): harden UDP session lifecycle (#286) * fix(traffic): harden UDP session lifecycle * fix(traffic): await UDP runtime shutdown * docs: prepare declarative apps release as v3 (#287) * docs(urls): update Gordon public links * docs(version): rename declarative apps release to v3 * docs(upgrade): clarify v3 migration guidance * refactor(apps): key manifest services by name (#288) * refactor(apps): key manifest services by name * fix(config): drop retired routes table from config example * fix(apps): restrict retired-shape hint to array tables * feat(apps): require --service or --all to deploy/restart multi-service apps (#292) * feat(apps): require --service or --all to deploy/restart multi-service apps App-wide deploy and restart of an app with several services now fail before any claim with ErrAppServiceScope (HTTP 400 service-scope-required), listing the sorted services. The admin API accepts 'all' (deploy body, restart query) and the CLI adds a mutually exclusive --all flag. apps apply --deploy keeps deploying every service. Closes #290 * fix(apps): scope check counts the revision to be deployed An unscoped deploy of an older revision slipped past requireServiceScope when the desired head and ACTIVE held a single service but the selected revision declared several: StartDeploy would then replace every service in that revision. Resolve the revision to activate first (explicit revision, else desired head) and count its services before claiming. Wrap the LoadDesired/LoadActive/LoadRevision scope reads with app context while preserving the underlying store errors. * feat(telemetry): export proxy access and app container logs over OTLP (#293) * refactor(telemetry): add Metrics port so registry and eventbus stop importing the adapter * feat(apps): add telemetry.logs opt-out at app and service level * feat(telemetry): export proxy access and app container logs over OTLP * fix(telemetry): export container output from collector start, not follower start * fix(telemetry): sanitize UTF-8, drop url.query, wrap state errors, persist resume cursors * fix(apps): skip unchanged services on deploy, apply secret changes on deploy (#294) * fix(apps): name the diverging path when a targeted deploy is refused * fix(apps): skip unchanged services on deploy and recreate containers on restart * fix(apps): keep deploy skip and restart safe for withdrawn services and missing secrets * fix(apps): apply secret changes on deploy and keep restart in place * fix(telemetry): export managed container gauge from v3 app state (#291) Replace the unused gordon.container.managed UpDownCounter with an observable gauge that counts ACTIVE app service containers (excluding stopped apps) at each collection. Closes #289. * chore(deps): bump Go dependencies (#295) * chore(deps): bump Go dependencies (grpc v1.84.0 fixes GHSA DoS alert) * chore(mocks): regenerate mocks with mockery v3.8 * refactor(cli)!: v3 command layout and remove domain secrets (#296) * refactor(cli)!: remove domain secrets, move push under images, add daemon group - Remove the v2 per-domain installation secrets stack (CLI, admin /secrets API, SecretService, domain secret store, env loader, env file migration, [env] config). No v3 code path read these values. - /admin/secrets now answers 410 endpoint-retired. - gordon push -> gordon images push. - gordon status/logs/reload/config/tls/traffic/networks -> gordon daemon <cmd>. - tls, traffic and networks are single commands under daemon. * docs: document daemon and images push commands, drop domain secrets - Add docs/cli/daemon.md covering status, logs, reload, config, tls, traffic, and networks; remove the per-command pages it replaces. - Fold push into docs/cli/images.md. - Remove gordon secrets, [env], env-variables, and admin:secrets docs; point users to gordon apps secrets. - Reject a leftover [env] config section with a config-retired diagnostic; drop dead env/container-name domain helpers. - Update README, guides, wiki, CI examples, and upgrade notes. * refactor: deepen daemon wiring, reload, app-state, deployment, and CLI modules (#297) * refactor(cli): let remote.Client implement ControlPlane directly Remove the remoteControlPlane pass-through wrapper. Both the explicit remote and the local admin socket already use *remote.Client, so the wrapper only forwarded calls and duplicated error context. * refactor(deployment): split candidate creation out of the deploy flow Move candidate container construction (env, volumes, image pull, create, journal, start) to candidate.go and return one startedCandidate value instead of four results. Backend bind handling, diagnostics, interrupted deploy convergence, and removed-service retirement move to their own files so execute.go reads as the deploy flow. * refactor(appstate): narrow app-state ports and hide apply sequencing Replace the stage/commit/materialize/GC call sequence in the apps use case with one AcceptApply method on the store, so the ordering invariant lives in the adapter. The apps use cases now depend on AppApplyStore and AppCatalogReader instead of the full AppState port. ListRevisions, LoadApplyIntent, and ListIntents leave the port; they had no use-case callers and stay on the concrete store for tests. * refactor(app): replace reload setter hooks with one ordered runtime step Move the reload coordinator to reload.go and replace the SetContainerConfigApplier and SetAppMountPoliciesApplier closures with a reloadRuntime passed at construction. It derives the container config and app bind/device policies before any side effect, so an invalid policy no longer leaves management hosts and traffic already updated. * refactor(app): split daemon wiring out of run.go by concern Move config loading, service construction, app engine wiring, auth and secrets, internal credentials, HTTP edge handlers, listeners, schedulers, shutdown, and PID file handling into dedicated files. run.go keeps Run and the server lifecycle (4213 -> ~380 lines). No behavior change. * fix(app): name the failing reload runtime step in errors * chore(app): drop orphaned nolint directive left by the run.go split * fix(app): build reload TLS config before any runtime side effect Load the static TLS keypair in reloadRuntime.prepare and rebuild the traffic graph before updating management hosts, so a failed keypair load or rejected graph leaves PKI and public TLS hosts unchanged. * fix(app): stop started schedulers when a later scheduler fails to start * fix(deployment): persist service removal before clearing its inhibition Save ACTIVE after each removed service and only then clear its recovery inhibition, so a failed save or later removal never leaves ACTIVE listing a service recovery is free to recreate. * fix(security): validate child manifest digests, cap log tail allocation, restrict CI token (#300)
Fixes the CodeQL alerts reported on #299 (
next→main).filesystem/manifest.go): child manifest digests from an image index were passed toGetManifestwithout format validation. They are now checked withvalidation.ValidateDigestbefore lookup (rejected asErrManifestBlobUnknown), andgetManifestPathadds an explicit..guard recognized by CodeQL. ExistingValidatePath+ root-containment checks already blocked traversal; this makes it explicit at the source.logs/service.go):tailLinescaps its ring buffer at 10000 lines inside the use case, not only in the HTTP handler.ci.yml): workflow-levelpermissions: contents: read.tokenstore/unsafe.go): false positive — SHA-256 hashes a token subject to build a safe filename, not a password. Dismissed in code scanning.Checks
go test ./...,golangci-lint run ./...: 0 issuesTestService_PutManifest_RejectsMalformedChildDigest