Operations
Tool / Contract Summary
This page describes the current operational posture of the live Loom runtime. It focuses on service health, readiness rules, monitoring expectations, and operator boundaries.
Current Verified State
Repo-owned runtime exists and is live on the owner-supplied runtime host.
Live stack members:
jhf-loom-dbjhf-loom-activemqjhf-loom-transformjhf-loom-searchjhf-loom-repojhf-loom-share
Canonical public path:
https://<internal-runtime-redacted>/
Available Now
Operational assets committed in-repo:
- compose.yml (
compose.yml) - compose.low-cpu.yml (
compose.low-cpu.yml) - compose.health-fast.yml (
compose.health-fast.yml) - config/runtime/runtime-contract.json (
config/runtime/runtime-contract.json) - config/runtime/alfresco-runtime-profile.json (
config/runtime/alfresco-runtime-profile.json) - config/runtime/alfresco-component-family.json (
config/runtime/alfresco-component-family.json)
The canonical Alfresco runtime profile is documented in:
- ALFRESCO_RUNTIME_PROFILE.md (
docs/ALFRESCO_RUNTIME_PROFILE.md)
Operational verify scripts:
validate_runtime_contract.pyvalidate_alfresco_runtime_profile.pyvalidate_alfresco_component_family.pyvalidate_runtime_materialization_drift.pyvalidate_readme_raw_truth.pyvalidate_live_proxy_sso_smoke.pyvalidate_live_runtime_resilience.pyvalidate_live_cold_start_recovery.pyvalidate_live_low_idle_runtime_policy.py
Operational control scripts:
loom-runtime-control.sh preflightloom-runtime-control.sh phased-startloom-runtime-control.sh applyloom-runtime-control.sh pauseloom-runtime-control.sh resumeloom-runtime-control.sh postcheckloom-runtime-control.sh diagnosticsloom-runtime-control.sh memory-guard
Upgrade-family inventory surfaces:
maintenance/pull_stack_oss_inventory.pyconfig/runtime/alfresco-component-family.json
Readiness / Drift / Monitoring
Hard runtime rules:
- search is mandatory for green readiness
- no partial-green state is allowed when search is red
- accepted failure posture is
degradedorhalted, never green - canonical Loom readiness gate is:
bash scripts/loom-runtime-control.sh readiness
- repository
/probes/-ready-is treated as repo-local signal only, not full Loom green readiness on its own - ActiveMQ is recoverable supporting infrastructure, not business truth
- shared-host startup must use the
shared-host-low-idlephased bring-up path - service activation posture is machine-readable in
SERVICE_ACTIVATION_MATRIX.md (
docs/SERVICE_ACTIVATION_MATRIX.md) and validated withpython scripts/validate_service_activation_matrix.py - default healthcheck cadence is
shared-host-slowfor shared-host baseline load control dedicated-fasthealthchecks are optional only and require explicitJHF_LOOM_HEALTHCHECK_PROFILE=dedicated-fast- heavy-profile startup on undersized hosts is blocked by preflight unless
JHF_LOOM_ALLOW_UNDERSIZED_HEAVY_PROFILE=1is explicitly set - default phased polling is low-pressure (
15s); tight watchdog loops are not a supported default - Phased startup cooldown spacing (
45s) is enabled by default to reduce warmup overlap for transform/repo/share applyskips services that are already healthy to avoid unnecessary restart pressure on shared hosts- compose/env drift is checksum-gated; changed inputs force rollout on the next
applyso config updates are not silently skipped - repo-owned control runs also hydrate missing non-secret runtime defaults from
.env.exampleinto the live stack.envbefore compose evaluation and update drifted non-secret budget keys when the repo-owned truth changes - non-default metadata keystore passwords require a matching host-managed
keystore artifact at
runtime-secrets/alfresco/extension/keystore/keystore; the drift validator must stay red until both the bind mount and artifact are present - runtime materialization drift must also be checked across repo files, live
stack files, container labels/env/mounts/ports/networks, and app readback:
python scripts/validate_runtime_materialization_drift.py --host <live-host> --stack-dir <owner-runtime-stack-dir> --insecure
- shared-host memory reclaim must be guarded before host-owner swap surgery:
bash scripts/loom-runtime-control.sh memory-guard- runbook: SHARED_HOST_MEMORY_RECLAIM.md (
docs/SHARED_HOST_MEMORY_RECLAIM.md)
loom-runtime-control.sh applyandloom-runtime-control.sh resumeenforce the same memory guard fail-closed before rollout/resume mutations- no-repeat lock prevents concurrent runtime-control mutation paths:
.loom-runtime-control/runtime-control.lock
- diagnostics must stay bounded (
timeout+--since+--tail) - restart recovery is bounded with backoff and optional jitter
- incident response may pause only non-critical services;
jhf-loom-dbremains the single kept-running service in pause mode - post-deploy cleanup runs after phased start/resume to remove stale stopped containers and rotate old diagnostic captures
- transform/repo/share run with Docker init reaping enabled (
init: true) to prevent zombie child accumulation during restart cycles
Current drift to track:
- public functional probes are green
- Docker can still report
jhf-loom-searchasunhealthydue to a stale health-taskNotFounderror on the host - this drift must not be documented as full operational perfection
Latest soak evidence:
- ALFRESCO_RUNTIME_SOAK_EVIDENCE_2026-04-21.md (
docs/ALFRESCO_RUNTIME_SOAK_EVIDENCE_2026-04-21.md) - ALFRESCO_HEALTHCHECK_PROFILE_EVIDENCE_2026-04-21.md (
docs/ALFRESCO_HEALTHCHECK_PROFILE_EVIDENCE_2026-04-21.md) - RUNTIME_GUARDRAILS_V1_EVIDENCE_2026-04-23.md (
docs/RUNTIME_GUARDRAILS_V1_EVIDENCE_2026-04-23.md)
Deployment / Verify
Repo:
python scripts/validate_repo_baseline.py
python scripts/validate_runtime_contract.py
python scripts/validate_alfresco_runtime_profile.py
python scripts/validate_alfresco_component_family.py
python scripts/validate_runtime_materialization_drift.py
python scripts/validate_readme_raw_truth.py
python scripts/validate_runtime_guardrails_v1.py
python scripts/validate_host_capacity_budget.py
python maintenance/pull_stack_oss_inventory.py --output test-results/stack-oss-inventory.workspace.json
python -m unittest discover -s tests -p "test_*.py"
Live:
python scripts/validate_runtime_contract.py --host <live-host>
python scripts/validate_alfresco_runtime_profile.py --host <live-host>
python scripts/validate_alfresco_component_family.py --host <live-host>
python scripts/validate_runtime_materialization_drift.py --host <live-host> --stack-dir <owner-runtime-stack-dir> --insecure
python maintenance/pull_stack_oss_inventory.py --host <live-host> --output test-results/stack-oss-inventory.workspace.json
python scripts/validate_live_low_idle_runtime_policy.py --insecure --idle-window-seconds 1800 --moderate-traffic-seconds 90
python scripts/validate_live_runtime_resilience.py --password <host-managed-secret> --insecure
python scripts/validate_live_cold_start_recovery.py --password <host-managed-secret> --insecure
bash scripts/loom-runtime-control.sh memory-guard
bash scripts/loom-runtime-control.sh diagnostics
bash scripts/loom-runtime-control.sh postcheck
bash scripts/loom-runtime-control.sh readiness
Scan&Fix Runbook
Schnellstart:
bash scripts/scan_and_fix.sh --dry-run
Gezieltes Issue:
bash scripts/scan_and_fix.sh --issue 89 --dry-run
Live-Run (Agent/CLI-Executor muss gesetzt sein):
export SCAN_FIX_EXECUTOR='codex exec --prompt-file {prompt_file}'
bash scripts/scan_and_fix.sh --host <live-host> --user <ssh-user>
Filter/Batch-Beispiele:
bash scripts/scan_and_fix.sh --labels runtime,bug --since 7d --max-issues 3 --dry-run
bash scripts/scan_and_fix.sh --severity-order "critical,high,medium,low" --dry-run
Fehlerbilder:
SCAN_FIX_EXECUTOR is required for non-dry-run execution:SCAN_FIX_EXECUTORsetzen oder--dry-runnutzen.
No matching open issues found:- Filter (
--labels,--since) lockern oder ohne Filter starten.
- Filter (
Unsupported origin URL format:- set
originto a supported Gitea remote format.
- set
- Live drift despite green repository checks:
- check the host path/stack and run the same verification path again after pushing.
Known Limits
- ingress, DNS, and TLS are not operated from this repo
- direct raw service ports are not canonical public paths
- host-level Docker health drift can diverge from the public-path functional posture and must be called out explicitly
License: AGPLv3.
Helpifyr: https://helpifyr.com
Workspace Git/Scan Guardrails (Mandatory)
- Gitea is Source of Truth; local Windows workspaces are disposable working copies.
- Never run Codex sessions on the workspace root; always use a concrete repo path.
- Limit active repo sessions to 2-3 in parallel.
- Before each run in a repo:
git fetch --prune,git checkout <branch>,git pull --ff-only. - No background git discovery loops (
git status,git ls-files, worktree scans) without explicit scoped need. - Automation scripts must run repo-scoped only, never global over the workspace root.
scan_and_fix Standard
scripts/scan_and_fix.shmust enforce runner timeout + single-run lock +.envfallback to the operator-managed workspace.env.
Release Eligibility Posture
- Public release-surface admission is bounded by
contracts/admission/release_surface_inventory_v1.json. - Current repository history posture is tracked in
contracts/admission/repository_publication_history_posture_v1.json. - Active owner blocker:
JaddaHelpifyr/jhf-loom#281. jhf-loomremainspending_history_scan; operator-only runtime runbooks keep live-host detail but are excluded from public release-surface admission.- Repo-owned history evidence now runs through
python scripts/verify_repository_history_posture.pyand currently provesfiltered_publication_required. scripts/scan_open_issues_repo_only.shmust exist and query only current repo open issues via Gitea API.
Workspace Hygiene
- Daily cleanup: stale
_worktrees/*,_tmp/*,test-results/*, large temporary artifacts. - Weekly cleanup: stale local branches/worktrees.
- Never leave valuable artifacts as untracked files in workspace root.
Dirty-State Policy
- Dirty state is allowed while actively implementing.
- Before new scan/automation runs: commit/stash, or use a dedicated worktree.
- Never propagate
dirty_unknownstates.
Incident Playbook (git.exe storm)
- Identify parent of
git.exe(usually oneCodex.exe). - Stop only the offending process tree.
- Restart session on concrete repo path.
- Reduce parallel sessions.
- Verify
git.execount drops within 30-60s.