Skip to content

fix(security): resolve CodeQL alerts blocking the v3 release - #300

Merged
bnema merged 1 commit into
nextfrom
fix/codeql-v3-release
Oct 1, 2026
Merged

bnema merged 1 commit into
nextfrom
fix/codeql-v3-release

Conversation

@bnema

@bnema bnema commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Fixes the CodeQL alerts reported on #299 (next → main).

  • go/path-injection (filesystem/manifest.go): child manifest digests from an image index were passed to GetManifest without format validation. They are now checked with validation.ValidateDigest before lookup (rejected as ErrManifestBlobUnknown), and getManifestPath adds an explicit .. guard recognized by CodeQL. Existing ValidatePath + root-containment checks already blocked traversal; this makes it explicit at the source.
  • go/uncontrolled-allocation-size (logs/service.go): tailLines caps its ring buffer at 10000 lines inside the use case, not only in the HTTP handler.
  • actions/missing-workflow-permissions (ci.yml): workflow-level permissions: contents: read.
  • go/weak-sensitive-data-hashing (tokenstore/unsafe.go): false positive — SHA-256 hashes a token subject to build a safe filename, not a password. Dismissed in code scanning.

Checks

  • go test ./..., golangci-lint run ./...: 0 issues
  • New test: TestService_PutManifest_RejectsMalformedChildDigest

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 46aa86ee-b25b-4218-b16d-b0395912d9cc

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@bnema
bnema marked this pull request as ready for review October 1, 2026 18:13
@bnema
bnema merged commit bb444bd into next Oct 1, 2026
5 checks passed
@bnema
bnema deleted the fix/codeql-v3-release branch October 1, 2026 18:14
bnema added a commit that referenced this pull request Oct 1, 2026
* feat: Gordon v2.50.0 application deployment model (#276)

* feat(apps): add declarative application deployments

* fix(security): enforce host trust boundaries

* fix(security): close remaining host trust gaps

* test(security): align hardened image fixtures

* refactor(cli): route local commands through admin socket

* fix(admin): authorize local control plane routes

* feat(install): build next channel from pinned source

* feat(install): default to user-local binary path

* fix(ci): avoid overlapping Go module caches

* fix(apps): address wrap-up review findings

* fix(apps): close v2.50 deferred findings (#280)

* fix(apps): persist mutation identity and close state cleanly

* refactor(cli): unify app control plane capabilities

* fix(apps): claim mutation keys atomically

Add a durable request identity to app operation journals and one atomic
store claim that binds an idempotency key to exactly one request:

- domain.AppOperation carries the immutable request identity (kind, app,
  requested revision, requested service).
- out.AppState.ClaimOperation persists the in-flight journal for an
  absent key, replays the stored journal for the same request, and
  refuses a key that already answered a different request with
  ErrAppStateConflict.
- A claim for a name with no live app identity fails with the new
  ErrAppNotFound and writes nothing; the admin API maps it to 404
  app-not-found. read-only AppExists backs the not-found normalization
  for revisions.
- Every keyed deploy/stop/start/restart/remove claims before any effect.
  A replayed key returns its stored journal and runs no workload call; an
  unfinished journal returns the stored record with a conflict.
- The apps use case returns the journaled operation together with the
  engine error instead of dropping the record.

* fix(apps): order operation steps behind publication and complete read models

Journal ordering:

- A deploy service step is recorded as succeeded only after the runtime
  effect, the ACTIVE publication, and the traffic graph apply succeed;
  a rejected apply fails the step instead of leaving a success behind.
- Lifecycle verbs record a failed traffic.publish step and a failed
  outcome when the final graph apply is rejected, so a rejected
  publication never replays as success.
- Every error path after the claim now records a terminal failure, so a
  keyed repeat replays the failure instead of reading a stuck claim.
- start marks a service step succeeded only after its effective state is
  published.

Read models:

- Show returns not-found for a name with no live app identity, and
  reports desired acceptance/pending status, per-service effective
  revision, container, digest, restart safety, owned resources
  (volumes, secret paths, image references), and the latest operation
  with kind, outcome, and start time.
- List skips retired names and reports desired status, pending,
  convergence, stop intent, and the last outcome. Pending is derived
  from desired vs ACTIVE, never from a journal.
- Admin DTO mappings carry these fields; deploy responses attach current
  effective and owned state instead of inferring convergence from the
  journal outcome. The unused observed field is gone.

* fix(app): publish the host index only after the graph apply succeeds

Traffic publication now prepares and commits instead of mutating the
live host index first:

- ACTIVE state projects into a detached candidate index that the graph
  is built and applied against. The live index is swapped only after the
  dataplane accepted the graph, so a rejected apply leaves both the
  index and the graph at the previous generation.
- A failing projection applies nothing at all.
- Withdrawal stays fail-closed: when the graph apply fails after the
  durable binds were cleared, the read model still drops the withdrawn
  host so the proxy never dials a recycled loopback port; the error is
  returned unchanged so the caller keeps its publication inhibition.

HostIndex gains an Entries accessor for the swap, and the publisher
takes its graph applier as a wired dependency so the ordering is
testable without a live traffic manager.

* fix(runtime): retire exact containers with the effective stop grace

Runtime contract:

- StopContainer and RestartContainer take the effective stop grace. The
  Docker adapter converts it to whole API seconds, rounding up, and keeps
  the runtime default for a non-positive grace instead of an immediate
  kill. The fixed twenty-second timeout is gone.
- Every app path passes the grace of the generation it stops: stop,
  restart, removal, replacement retirement, readiness-failure cleanup,
  and recovery cleanup. A missing declared grace falls back to the app
  default, never to zero.

One retirement path:

- retireContainer stops with grace (or force-removes a never-published
  candidate), confirms the exact ID is gone, and only then releases the
  container's backend claims and clears its recovery inhibition. A
  container that is not confirmed gone keeps its claims, its inhibition,
  and its ACTIVE reference, and fails the operation.
- Stop, remove, deploy service removal, HTTP replacement retirement,
  interrupted replacement, readiness-failure candidate cleanup, create
  and start cleanup, backend-bind registration failure, and
  container-not-found recovery all use it, so no disappearance path leaks
  a loopback claim.
- A confirmed-missing container releases its claims during recovery.

Bounded leftovers:

- Journaled operations carry bounded cleanup warnings (service, leftover,
  detail) for anything that could not be removed or released; they reach
  the admin response instead of being dropped, and never contain logs.

* fix(apps): reclaim private app networks on removal

Removal now removes the app's incarnation-owned private networks
immediately, once every exact container is confirmed gone and while the
ownership record is still live:

- The runtime name is re-derived from the incarnation UUID and the
  observed labels must prove that exact ownership, so a name reused by a
  later incarnation can never delete the previous one's network.
- A network with attached containers, or with labels that do not prove
  this incarnation, is left in place and reported as a bounded warning.
- A failed removal aborts the operation before the incarnation is
  retired, so the stopped intent, the inhibition, and ACTIVE state stay
  and a retry can converge.
- Shared networks are never removed by an app removal, and retained
  volumes, secrets, and images are untouched.

* refactor(backup): resolve app backups from manifest declarations

Backup targets are now read from the app's ACTIVE record instead of being
inferred, and the app name is the backup identity:

- A target exists because a service declares it: a PostgreSQL database in
  service.databases referenced by service.backup.postgres, or a volume in
  service.volumes referenced by service.backup.volume. The image, port,
  container label, and attachment heuristics are gone, along with the
  detection endpoint and DTOs that exposed them.
- Selectors are explicit: an omitted --service or resource selector only
  succeeds when exactly one compatible target exists, ambiguity lists the
  candidate service/resource names, and an unknown target is not found.
- The scheduler runs only databases whose own declared schedule matches
  the firing tier, and volume runs follow the installation preset.
- Artifacts are stored under the app name: filesystem paths and S3 keys
  use apps/<app>/..., never a domain prefix. Existing domain-keyed
  artifacts are left untouched and are not read.
- The admin API and CLI rename domain fields to app, add service, and
  name the resource database or volume. The CLI gains
  `backup run APP --service S --database D` and
  `backup volume run APP --service S --volume V`; the domain-keyed
  `detect` command and the databases/volumes aliases are gone.

* refactor(cli): unify the app and daemon control planes

The CLI now has one ControlPlane interface and one daemon-backed
implementation:

- App methods moved into ControlPlane; AppControlPlane,
  remoteAppControlPlane, and app_controlplane.go are gone, and
  remoteControlPlane delegates every operation to a single remote.Client
  (explicit remote or owner-only local admin socket).
- App commands resolve through resolveAppPlane, which returns a handle
  they close, instead of a separate resolver with no ownership.
- CLI tests drive the generated MockControlPlane: the handwritten
  fakeAppPlane and status shim are gone, and the source-text parity and
  UI adoption checks are replaced by the behavior they guarded (local
  socket and explicit-remote resolution, DTO parity across transports,
  presentation helpers).

* refactor: drop dead app plumbing and test-only production hooks

- AppServiceImpl keeps one configured core apps Service instead of
  rebuilding it on every apply; the listener, barrier, and image-policy
  options configure that instance in place.
- Remove the container cache/drain/metrics hooks that only stored their
  dependencies, the proxy container-deployed invalidation handler and its
  event payload, the container event constants nothing published or read,
  the internal-deploy context flag with no reader, the label constants
  the declarative cutover stopped using, and the deprecated revision
  retention alias.
- Move the app-state corruption hooks out of the production store into
  export_test.go.

* docs: describe current backup, log, and app identity behavior

- Backup docs document declarative targets (app, service, database or
  volume), the explicit selector rules, app-keyed artifact paths, and the
  app-based JSON fields.
- App docs describe the idempotency contract, the complete show read
  model, and app/service-only log identity.
- `gordon logs` reads daemon process logs only: the domain-keyed
  container-log path and its control-plane methods are gone, and app
  workload output is read through `gordon apps logs`.

* fix(apps): close deferred review gaps

* feat(apps): fix migration friction and add internal HTTP interfaces (#282)

* fix(apps): reduce migration friction

* feat(apps): add internal HTTP interfaces

* fix(apps): harden internal deployment lifecycle

* fix(apps): preserve targeted deployment state (#283)

* fix(apps): preserve networks in active diffs

* fix(apps): allow safe targeted deployments

* fix(apps): preserve non-targeted services

* fix(apps): parse shared network tables

* refactor(deployment): use durable sequential app deploys (#284)

* refactor(deployment): use sequential service replacement

Replace the rolling-versus-interrupted split for app services with one
sequential replacement applied to every service:

  preflight -> withdraw traffic (confirmed) -> retire the superseded
  container (confirmed gone) -> create and start the replacement ->
  readiness probe -> publish ACTIVE -> rebuild traffic

A deploy may briefly interrupt a service, and two Gordon-managed
generations of one service never run at the same time.

* refactor(deployment): drop httpEligible, deployHTTP, deployInterrupted,
  cutoverService, retireContainerID, drainWithDeadline and the
  retire-after-publish path
* refactor(deployment): restart always restarts the same pinned container
  in place (withdraw, restart, readiness, republish) and no longer creates
  a second container
* refactor(deployment): resolve the implicit HTTP readiness probe from the
  service spec so readiness no longer depends on how a service is
  replaced, keeping the check for stateless public HTTP services
* refactor(deployment): remove DeployResult.Interrupted and
  ServiceResult.Retire/RetireGrace
* test(deployment): replace the rolling-behaviour tests with sequential
  replacement, in-place restart, readiness-ordering, withdrawal-failure,
  cancellation, recovery-inhibition and multi-service failure tests
* docs: describe sequential replacement, the expected short interruption,
  preflight protection and manifest-based rollback

* fix(deployment): recover interrupted replacements safely

Close the crash window left by sequential replacement and the recovery
dead ends the review of it exposed.

* fix(deployment): journal the replacement's container ID before it is
  started, and keep it in the journal when a replacement fails after
  creating its container, so a leftover stays traceable
* fix(deployment): reconcile an unfinished operation at boot and before
  every mutation of an app: remove a never-published candidate (releasing
  its loopback claims) before creating or rebuilding a generation, keep
  the published generation, retry the leftovers of an operation that
  already failed, and finalize the operation (failure, or success when
  every step had already completed) so a keyed retry never reads a
  permanent in-flight claim
* fix(deployment): refuse a new journal while a different operation of the
  same app is unfinished, so a claim can never mask an interrupted
  predecessor
* fix(deployment): a restart that rebuilds a missing generation clears the
  recovery inhibition it wrote for the container proven gone, which
  otherwise refused every later start, recovery pass, and restart
* refactor(deployment): recover the store before reconciling in deploy,
  move the preflight store recovery to its callers, and extract
  runtimeImageRef, convergeCandidate, and interruptedOutcome
* refactor(config): remove the dead [deploy] configuration surface
  (pull_policy, readiness_*, health_timeout, stabilization_delay, probe
  timeouts, drain_*) from the viper defaults, the example configuration,
  and the docs: no code ever read these keys
* test(deployment): cover the candidate journal write before start, boot
  and mutation-time reconciliation, terminal-operation leftovers, an
  interruption after the last step, the claim guard, restart rebuilding a
  missing container, and the inhibition cleanup
* docs: document crash recovery, the 15 s republish cadence, and restart
  rebuilding from the pinned digest

* fix(deployment): fail closed on orphan replacements and repair recovery

Close the recovery findings left by the sequential-replacement review.

* fix(deployment): fail closed when a replacement candidate cannot be removed.
  Reconciliation returns an error, keeps the journal (and its latest position)
  instead of finalizing it, and no later mutation proceeds, so a newer
  operation can never mask an orphan that may still run.
* fix(deployment): clear the superseded container's stale replacement-pending
  inhibition once its never-published candidate is confirmed gone, so
  boot/start/restart can rebuild the generation ACTIVE records. The inhibition
  is kept while candidate cleanup fails, preserving the single writer.
* fix(deployment): set Service on start/restart placeholder journal steps so a
  recovered step is attributable and its inhibition can be cleared.
* fix(deployment): repair the Preflight(key)->Deploy(same key) contract:
  Preflight reconciles before resolving and never claims the caller's request
  key, so a later Deploy of the same key executes instead of replaying the
  preflight journal as a conflict.
* refactor(deployment): drop the unused bool result of verifyRunningService.
* test(deployment): cover fail-closed candidate removal, stale inhibition
  removal with a volume-owning rebuild, the repaired preflight/deploy key
  contract, and preflight reconciliation of an unfinished predecessor.
* docs: describe fail-closed reconciliation and inhibition removal.

* refactor(deployment): split deploy engine into start and execute phases

Factor the synchronous deploy engine into a durable two-phase internal
API. StartDeploy runs under the app lock: it reconciles any interrupted
predecessor, resolves the revision and runs the convergence checks,
atomically claims the request key, and persists a non-terminal journal.
It performs no image pull, pinning, or workload mutation, and reports
whether the caller newly owns the operation or is an idempotent replay.

ExecuteDeploy reacquires the app lock, loads and verifies the claimed
journal by app/op identity, runs image pull/pinning and the existing
sequential replacement, and records the terminal outcome. A terminal or
mismatched claim is never executed: it replays instead. Cancellation or
shutdown mid-execution leaves the non-terminal journal for reconciliation.

Service.Deploy stays the synchronous composition of both phases, so all
existing callers and behavior are unchanged. Existing preflight is
re-expressed as claimDeploymentLocked plus pinPreflightLocked.

* feat(apps): orchestrate two-phase deploy in the background

AppServiceImpl.Deploy now claims and persists the operation through
StartDeploy and, for the key owner only, schedules ExecuteDeploy on a
daemon-owned context so the request returns the running journal promptly
and request cancellation never aborts the replacement. Duplicate keys
replay their stored journal and never spawn a second execution.

A daemon context plus WaitGroup on AppServiceImpl gives the lifecycle
its graceful shutdown hook: Shutdown cancels in-flight executions,
which leave the persisted non-terminal journal for reconciliation, and
waits for them to unwind.

* feat(app): wire daemon lifecycle context into app administration

Build AppServiceImpl through newAppDaemonService with the
daemon/supervisor lifecycle context, so background deploy executions are
owned by the daemon and never by a request context.

Graceful shutdown registers AppServiceImpl.Shutdown on the bounded 30s
shutdown context as Phase 0.5, cancelling and joining in-flight deploy
executions before traffic, runtime, or the app state store is torn down.
A failed unwind is logged and never aborts teardown.

* feat(apps): answer 202 for running deploys and quiesce kernel app admin

The admin deploy handler now answers 202 Accepted with the durable running journal when it newly owns the operation, so the request returns promptly and the client observes completion via GET operations/by-key. A terminal replay stays 200, and conflicts, preflight failures, and failed replays keep their existing 409 mapping. AppDeployResponse carries a stable status field (running until terminal, then the persisted outcome); the by-key lookup keeps answering 200 and remains the recovery source. Permission checks and Idempotency-Key validation are unchanged, and the remote client's generic 2xx path already accepts 202.

Kernel.Close now cancels and joins the daemon-owned AppServiceImpl on a bounded context before any remaining kernel resource is torn down, so a background deploy execution cannot outlive the state and runtime it uses. A nil implementation is never stored in the lifecycle interface.

* feat(apps): poll 202 deploy operations and add operations watch

`gordon apps deploy` now follows a backend 202 through the existing
operations/by-key endpoint until the operation is terminal, printing
concise progress only when operation or step state changes. JSON mode
keeps stdout to a single final document (progress goes to stderr) and
terminal partial/failed outcomes exit nonzero.

Add `gordon apps operations watch APP --key KEY` to resume polling an
existing operation without reissuing the mutation. Ctrl-C stops only
local polling, never the daemon-side operation, and prints the resumable
command in human mode. 404 fails fast; transient polling errors are
bounded; local and remote planes share the same client path.

Docs: apps deploy/operations sections and CLI index.

* fix(deployment): keep live deploy claims out of foreground reconciliation

Track the operations this process owns between the StartDeploy claim and the
ExecuteDeploy exit, so a foreground mutation landing in the hand-off window
cannot reconcile and terminalize a live claim. ExecuteDeploy now always
reloads the operation from the store by app/op identity and refuses a
terminal or mismatched record, never executing a stale in-process journal.

* fix(apps): settle deploy claims raced by shutdown before scheduling

Gate AppServiceImpl.Deploy on the daemon lifecycle before StartDeploy, so a
request arriving after shutdown began is refused with a wrapped
domain.ErrAppStateConflict and never claims a journal or answers 202. Enroll
the Deploy hand-off with the lifecycle so Shutdown waits for an in-flight claim
to be either scheduled or settled, closing the window where shutdown starts
after StartDeploy but before scheduleExecution.

When shutdown wins that race, settle the owned claim instead of dropping it:
add deployment.Service.AbandonDeploy, which releases the in-process live marker
and converges the still non-terminal journal as interrupted, so no accepted
durable claim stays marked live forever and a later claim recovers cleanly.

Add deterministic tests for stopping before a deploy and for shutdown racing
between the claim and scheduling, plus an engine test proving the settled claim
is terminal and never blocks the next owner.

* fix(apps): reuse deploy watch for accepted apply --deploy

`apps apply --deploy` now follows a 202 accepted deploy through the same
by-key watch as `apps deploy`: progress only when the operation or step
state changes, one terminal journal, and a nonzero exit for a terminal
partial/failed outcome without hiding the successful apply.

JSON mode emits exactly one final combined {apply,deploy} document and
sends progress, transient warnings, and Ctrl-C resume guidance to stderr;
no initial running document is written and the deploy is never re-POSTed.
The command now threads cmd.ErrOrStderr so stderr routing is exercised.

Add focused tests for 202 to success, partial/failure with preserved
apply, unchanged-state deduplication, JSON purity, single mutation, and
Ctrl-C resume in both human and JSON modes.

* fix(apps): return journal shape for running deploy replays

A same-key deploy replay whose operation is still in flight now always answers 409 with the stored operation journal, before or after any service step has started. Only an absent or terminal journal may map a preflight error envelope, so a non-terminal replay never degrades to the mapped error response.

* refactor(deployment): remove dead exported Service.Preflight

Remove the production-dead deployment.Service.Preflight entry point and
its now-unreachable preflightLocked wrapper. No production caller and no
boundary interface exposed it; only direct tests did. The private
preflight gates (claimDeploymentLocked/pinPreflightLocked) used by
StartDeploy/ExecuteDeploy are preserved.

Delete the tests and helpers that existed only for the removed entry
point.

* test(deployment): restore preflight gate coverage through deploy phases

Retarget the live preflight gates that were only reachable through the removed Service.Preflight onto StartDeploy/ExecuteDeploy and the synchronous Deploy composition: missing required secret (ErrAppSecretMissing), env/secret collision (ErrInvalidAppSpec), unmanaged image volume (ErrAppUnmanagedImageVolume), R2 unowned existing volume refusal (ErrAppStateConflict), the targeted-deploy diff policy, and the shared GC barrier held across resource selection.

Every failing gate proves no workload mutation before the failure, and the execute-phase gates prove a terminal failed journal. No production code changes.

* fix(deployment): address HTTP readiness at declared container port

qBittorrent rejects the readiness probe with 401 "Invalid Host header, port mismatch" because the probe dials the random loopback published port and Go derives the Host header from it. Keep dialing the published bind but send 127.0.0.1:<declared container readiness port> as the Host authority, matching what the workload listens on.

Thread the Host authority through the injectable httpGet dependency so tests can assert both the dial URL and the Host header, and cover the mismatch with a server that rejects published-port authorities.

* fix(deployment): harden sequential deploy lifecycle

Address the review findings on the sequential service replacement work:

- keep the revision resolved at claim time authoritative for the whole
  execution instead of re-resolving desired state in the second phase
- record successful candidate convergence by clearing the step After and
  keeping the container id in Detail, so a converged step is not retried
  while its audit trail survives
- normalize the persisted service-step plan when a claim is resumed and
  converge recorded candidates before the plan is replaced
- fail closed when app administration does not quiesce: skip state and
  runtime teardown and propagate the error to the process entry point
- carry app and op fields through zerowrap contexts for the deploy
  scheduling logs instead of the service logger
- move the operations watch command into apps_operations_watch.go
- exercise the new admin deploy handler cases over a real httptest server

* feat(apps): add allowlisted CDI devices for app services (#285)

* feat(apps): add allowlisted CDI devices for app services

Manifests declare logical devices via devices = [...]; the
installation maps each name to explicit CDI IDs plus exact
app/service allowlists under [app_devices.<name>]. Resolution
happens at activation time and encodes one native CDI DeviceRequest;
revisions persist logical names only. Revoked grants and unsupported
engines fail closed before any workload mutation.

* refactor(apps): extract helpers to satisfy gocyclo budget

Split device diff entry, preflight error tail, and reload policy
validation out of diffService, isMappedPreflightError, and
applyLoadedConfig.

* fix(apps): address device review findings

Probe engine CDI capability in preflight via a new
ContainerRuntime.SupportsCDIDevices boundary method so unsupported
engines fail before workload mutation. Preserve ErrRuntimeUnsupported
through createContainer redaction and map it to the structured
runtime-unsupported envelope. Stop echoing resolved CDI IDs in
duplicate-grant errors and run union resolution at apply time.

* fix(apps): harden device error sanitization

Hoist the runtime-unsupported sentinel above the bind redaction
branch so mixed bind+device creates keep the structured envelope.
Sanitize the preflight engine probe cause so daemon endpoints never
reach journals or API responses.

* fix(apps): harden device policy, engine gate, and probe errors

Address review findings on the CDI device work:

- docs: move `devices` above the first nested service table so TOML
  assigns it to the service, not to the last bind entry.
- domain: keep the ErrInvalidAppSpec chain through device policy
  allowlist validation, and reject CDI device names that are host paths
  or raw device nodes instead of valid CDI names.
- docker: parse engine major/minor strictly, so prerelease suffixes such
  as "-rc1" fail closed instead of clearing the version gate; keep
  cancellation and deadlines from the /version probe while dropping the
  cause text that embeds the daemon endpoint.
- deployment: preserve cancellation and deadlines from the capability
  probe and collapse every other cause to ErrRuntimeUnsupported alone.
- appstate: load the legacy no-devices record from a verbatim fixture
  instead of writing through the current serializer.

* fix(traffic): harden UDP session lifecycle (#286)

* fix(traffic): harden UDP session lifecycle

* fix(traffic): await UDP runtime shutdown

* docs: prepare declarative apps release as v3 (#287)

* docs(urls): update Gordon public links

* docs(version): rename declarative apps release to v3

* docs(upgrade): clarify v3 migration guidance

* refactor(apps): key manifest services by name (#288)

* refactor(apps): key manifest services by name

* fix(config): drop retired routes table from config example

* fix(apps): restrict retired-shape hint to array tables

* feat(apps): require --service or --all to deploy/restart multi-service apps (#292)

* feat(apps): require --service or --all to deploy/restart multi-service apps

App-wide deploy and restart of an app with several services now fail before any claim with ErrAppServiceScope (HTTP 400 service-scope-required), listing the sorted services. The admin API accepts 'all' (deploy body, restart query) and the CLI adds a mutually exclusive --all flag. apps apply --deploy keeps deploying every service.

Closes #290

* fix(apps): scope check counts the revision to be deployed

An unscoped deploy of an older revision slipped past requireServiceScope
when the desired head and ACTIVE held a single service but the selected
revision declared several: StartDeploy would then replace every service
in that revision. Resolve the revision to activate first (explicit
revision, else desired head) and count its services before claiming.
Wrap the LoadDesired/LoadActive/LoadRevision scope reads with app
context while preserving the underlying store errors.

* feat(telemetry): export proxy access and app container logs over OTLP (#293)

* refactor(telemetry): add Metrics port so registry and eventbus stop importing the adapter

* feat(apps): add telemetry.logs opt-out at app and service level

* feat(telemetry): export proxy access and app container logs over OTLP

* fix(telemetry): export container output from collector start, not follower start

* fix(telemetry): sanitize UTF-8, drop url.query, wrap state errors, persist resume cursors

* fix(apps): skip unchanged services on deploy, apply secret changes on deploy (#294)

* fix(apps): name the diverging path when a targeted deploy is refused

* fix(apps): skip unchanged services on deploy and recreate containers on restart

* fix(apps): keep deploy skip and restart safe for withdrawn services and missing secrets

* fix(apps): apply secret changes on deploy and keep restart in place

* fix(telemetry): export managed container gauge from v3 app state (#291)

Replace the unused gordon.container.managed UpDownCounter with an observable gauge that counts ACTIVE app service containers (excluding stopped apps) at each collection. Closes #289.

* chore(deps): bump Go dependencies (#295)

* chore(deps): bump Go dependencies (grpc v1.84.0 fixes GHSA DoS alert)

* chore(mocks): regenerate mocks with mockery v3.8

* refactor(cli)!: v3 command layout and remove domain secrets (#296)

* refactor(cli)!: remove domain secrets, move push under images, add daemon group

- Remove the v2 per-domain installation secrets stack (CLI, admin
  /secrets API, SecretService, domain secret store, env loader, env file
  migration, [env] config). No v3 code path read these values.
- /admin/secrets now answers 410 endpoint-retired.
- gordon push -> gordon images push.
- gordon status/logs/reload/config/tls/traffic/networks -> gordon daemon <cmd>.
- tls, traffic and networks are single commands under daemon.

* docs: document daemon and images push commands, drop domain secrets

- Add docs/cli/daemon.md covering status, logs, reload, config, tls,
  traffic, and networks; remove the per-command pages it replaces.
- Fold push into docs/cli/images.md.
- Remove gordon secrets, [env], env-variables, and admin:secrets docs;
  point users to gordon apps secrets.
- Reject a leftover [env] config section with a config-retired
  diagnostic; drop dead env/container-name domain helpers.
- Update README, guides, wiki, CI examples, and upgrade notes.

* refactor: deepen daemon wiring, reload, app-state, deployment, and CLI modules (#297)

* refactor(cli): let remote.Client implement ControlPlane directly

Remove the remoteControlPlane pass-through wrapper. Both the explicit
remote and the local admin socket already use *remote.Client, so the
wrapper only forwarded calls and duplicated error context.

* refactor(deployment): split candidate creation out of the deploy flow

Move candidate container construction (env, volumes, image pull, create,
journal, start) to candidate.go and return one startedCandidate value
instead of four results. Backend bind handling, diagnostics, interrupted
deploy convergence, and removed-service retirement move to their own
files so execute.go reads as the deploy flow.

* refactor(appstate): narrow app-state ports and hide apply sequencing

Replace the stage/commit/materialize/GC call sequence in the apps use
case with one AcceptApply method on the store, so the ordering invariant
lives in the adapter. The apps use cases now depend on AppApplyStore and
AppCatalogReader instead of the full AppState port. ListRevisions,
LoadApplyIntent, and ListIntents leave the port; they had no use-case
callers and stay on the concrete store for tests.

* refactor(app): replace reload setter hooks with one ordered runtime step

Move the reload coordinator to reload.go and replace the
SetContainerConfigApplier and SetAppMountPoliciesApplier closures with a
reloadRuntime passed at construction. It derives the container config
and app bind/device policies before any side effect, so an invalid
policy no longer leaves management hosts and traffic already updated.

* refactor(app): split daemon wiring out of run.go by concern

Move config loading, service construction, app engine wiring, auth and
secrets, internal credentials, HTTP edge handlers, listeners, schedulers,
shutdown, and PID file handling into dedicated files. run.go keeps Run
and the server lifecycle (4213 -> ~380 lines). No behavior change.

* fix(app): name the failing reload runtime step in errors

* chore(app): drop orphaned nolint directive left by the run.go split

* fix(app): build reload TLS config before any runtime side effect

Load the static TLS keypair in reloadRuntime.prepare and rebuild the
traffic graph before updating management hosts, so a failed keypair
load or rejected graph leaves PKI and public TLS hosts unchanged.

* fix(app): stop started schedulers when a later scheduler fails to start

* fix(deployment): persist service removal before clearing its inhibition

Save ACTIVE after each removed service and only then clear its recovery
inhibition, so a failed save or later removal never leaves ACTIVE
listing a service recovery is free to recreate.

* fix(security): validate child manifest digests, cap log tail allocation, restrict CI token (#300)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant