Z-Arch Platform Documentation
Table of Contents
I. Understanding Z-Arch
Introduction
Why Z-Arch exists
Application logic is cheap and easy to generate. Backend security architecture is not. A solo developer can produce routes, handlers, and database calls in a day. What they cannot reliably produce with the same momentum is a correct perimeter, a consistent authentication model, and defense-in-depth in a scalable shape that does not leak cost or drift over time. The usual result is either improvised security or reliance on hosted platforms that own your infrastructure, data, and wallet.
Z-Arch closes the gap cleanly by defining a fixed backend architecture in one small configuration file, zarch.yaml. Authentication, perimeter enforcement, and service trust are enforced automatically on serverless cloud infrastructure you own and control. This is not another AI code generator or agent skill. Nor is it a “codebase security scanner” that looks for problems only after they exist. It is proactive, deterministic software that constrains and guards your application code in the simplest manner possible. You describe your system in zarch.yaml. Z-Arch makes that description safe and guaranteed.
In practice this means:
All public traffic passes through a single enforced ingress plane.
Backend services run on an internal private network; the Gateway service is the only backend runtime intended to accept outside traffic directly.
No request is trusted, not even internally.
Requests that reach your business logic have already satisfied the authentication guarantees defined by the platform.
Scale-to-zero is the default economic behavior of every component.
Your entire backend architecture is described in an elegant and concise format.
With the architecture guaranteed, a single developer, or an LLM operating within the constraints of the Z-Arch MCP, can generate a complete and secure backend with ease. Business logic becomes the cognitive focus. Your AI saves context, you save peace of mind.
The Z-Arch Philosophy
Z-Arch is Post-SaaS software. It does not host your application. It does not store your data. It does not introduce a proprietary runtime you cannot leave. Everything deploys into your own GCP project. You own the infrastructure, the secrets, the code, and the data. The architecture is opinionated by design, because the goal is correctness, repeatability, and security at minimal cost. But unlike alternative services, it never stands in your way. Z-Arch is anti lock-in.
Platform Architecture
When you bootstrap a project with Z-Arch, you will see a small set of platform-owned resources appear alongside your own services and jobs. These are the stable components Z-Arch uses to apply zarch.yaml, enforce the API perimeter, and keep your project convergent over time.
At a high level, Z-Arch separates the system into three concerns: deployment orchestration, the API perimeter, and your application workloads. Each concern has a dedicated component, and those components keep their responsibilities narrow.
Z-Arch Managed Components
Control Plane
Z-Arch deploys a Control Plane job into your project. This is the execution surface that applies zarch.yaml and performs platform operations in a defined order. It is not part of your runtime request path. It exists so infrastructure and policy are applied deterministically rather than by ad-hoc scripts or manual console work.
The Control Plane and Extension Runner clone the connected repository through a project-specific, read-only SSH deploy key. Z-Arch stores the private half in the project’s Secret Manager. The repository’s owning organization or enterprise must allow deploy keys; see GitHub Prerequisites.
Gateway
Z-Arch deploys the Gateway as a Cloud Run service. This is the single API perimeter for your backend and the only backend runtime intended to accept outside traffic. Requests are routed and authenticated here before reaching your services. Your services are not public ingress points.
Extension Runner
You may also see an Extension Runner job. Its presence is normal. It exists to execute configured extension lifecycle logic as part of Z-Arch operations.
Where Your Code Runs
Edge Proxy
Z-Arch deploys a Cloudflare Worker, the Edge Proxy, in front of the Gateway. The default edge resolves API keys, rewrites the edge-only client path, selects a configured Gateway, attaches Google OIDC, and benefits from Cloudflare’s network-level DDoS protection. The edge proxy source is completely customizable, offering robust opportunity for implementing custom logic before a request reaches the gateway.
The edge proxy is not where Z-Arch’s managed authentication model is enforced. The Gateway is.
Services
Your services are Cloud Run services that implement application logic. They sit behind the Gateway on an internal private network. In the intended topology, they only accept internal ingress on a private network, receive requests that have already passed gateway enforcement, and do not accept outside traffic directly.
On GCP, Z-Arch manages the VPC network, regional subnet attachment, Private Google Access, and private DNS configuration so backend Cloud Run traffic stays on the private path.
Jobs
Your jobs are Cloud Run jobs for run-to-completion workloads that don’t need to serve HTTP traffic. They are invoked by schedulers, event flows, or directly by services and can interact with other resources according to declared targets.
Topics and Schedulers
Topics and schedulers are the event and time primitives that connect services and jobs without turning everything into direct synchronous calls. They are declared in zarch.yaml and deployed as part of the same convergent system.
Request Flow
Public ingress in the default topology is:
Client -> Edge Proxy -> Gateway -> Service
The Gateway -> Service hop is an internal backend hop.
The operational boundary sequence is:
- Edge proxy validates client API-key posture and forwards traffic to the gateway.
- Gateway matches route and enforces declared auth mode (
dual,jwt-only,cookie-only, orpublic). - For protected routes, the gateway completes the JWT, session-cookie, and CSRF checks required by that route before proxying.
- Gateway attaches short-lived internal trust assertions to the upstream call.
- Target service verifies S2S trust assertions before accepting the internal call.
Service-to-service calls follow the same trust model: callers mint short-lived assertions and receivers verify caller identity, intended audience, and token freshness.
Container boundary rule: protected traffic does not execute inside an application service container until perimeter checks have passed at the gateway.
Gateway Enforcement Model
The Gateway is the single API perimeter for a Z-Arch system. It exists so authentication and perimeter controls are enforced before any request reaches your application container.
Application services are not ingress points. The Gateway is.
Routing and Request Matching
The Gateway is the only component responsible for mapping incoming requests to backend services.
- Routes are declared in
zarch.yaml. - Path and method evaluation are deterministic.
- Literal and regex route styles are supported.
- Endpoint exposure is controlled per service.
If a route is not declared, it does not exist.
JWT Verification with OIDC
For routes that require a JWT (dual and jwt-only), the Gateway validates end-user JWTs against your configured OIDC provider.
- OIDC discovery metadata is configurable.
- JWT signature and audience checks are enforced.
- Validation is based on OIDC/JWKS standards.
Services do not parse or validate end-user JWTs for public traffic or for cookie-only routes. When a route requires JWT validation, that enforcement occurs at the Gateway.
Session Enforcement and CSRF Protection
For protected routes using the dual or cookie-only auth modes, Z-Arch enforces a valid session cookie before traffic reaches the target service.
For dual routes, Z-Arch requires:
- A valid end-user JWT.
- A valid session cookie bound to that JWT.
For cookie-only routes, Z-Arch requires:
- A valid session cookie.
Routes using dual or cookie-only automatically enable CSRF protection using the double submit cookie method:
- The gateway returns a non-HttpOnly
X-CSRF-Tokencookie. - The frontend client must send that same value back in the
X-CSRF-Tokenrequest header. GET,HEAD, andOPTIONSare treated as safe methods and are not subject to CSRF validation.- Other methods are unsafe and must present a matching cookie/header pair.
- A fresh CSRF cookie is returned on each request.
- Each CSRF cookie expires after
7200seconds.
This enforcement occurs before traffic is forwarded to the target service. A request never executes inside the same container as your application logic unless it satisfies its declared auth mode.
Route Auth Modes
Auth mode is declared per route.
Available modes:
dual— JWT and session cookie required.jwt-only— JWT required, no session cookie.cookie-only— session cookie required, no JWT.public— no end-user authentication required.
Routes using dual or cookie-only also automatically enable CSRF protection via the double submit cookie method.
Even public routes remain inside the Gateway perimeter and are not directly exposed as independent services.
CORS and Security Headers
CORS policy and security headers are enforced at the Gateway level.
- Allowed origins, headers, and methods are centrally configured.
- Credential mode and cache windows are applied uniformly.
- Security headers are applied consistently across routes.
- Browser clients calling
dualorcookie-onlyroutes should allow theX-CSRF-Tokenrequest header through CORS.
Services do not manage CORS logic. This prevents inconsistent browser-facing behavior across services.
Request Validation and Rate Controls
The Gateway enforces route-level operational controls:
- Required content type enforcement.
- Per-route rate limit declarations.
- Capability checks where
authzis configured.
These controls are applied before forwarding to services. Services do not implement their own rate limiting for public traffic unless explicitly required for domain reasons.
Internal Service Trust Propagation
When the Gateway forwards a request to a service, it attaches short-lived internal trust token.
Services verify this assertion with ZArchAuth.s2s.verify() to ensure the caller is the Gateway and that the request is intended for them.
What You Still Implement in Your Services
The Gateway does not make domain-level authorization decisions.
Services remain responsible for:
- Business authorization rules.
- Domain-specific input validation.
- Application error handling and observability.
- Data-layer constraints and integrity.
The Gateway enforces perimeter authentication and route policy. Application services enforce business logic.
Authentication and Access Model
OIDC Model
Z-Arch validates end-user JWTs using any OIDC-compliant provider.
Primary inputs:
authn.discovery: OIDC discovery URL.authn.client_id: expected JWT audience.
Managed Firebase bootstrap is available under authn.firebase, but runtime JWT validation is provider-agnostic under the same OIDC model.
Routes using cookie-only rely on the established session cookie for that request and do not require an end-user JWT on the route itself.
JWT Validation, Session Binding, and CSRF
For authenticated routes, gateway policy enforces the configured route auth mode.
dual: both a valid JWT and a valid session cookie bound to that identity are required.jwt-only: a valid JWT is required; no session cookie requirement.cookie-only: a valid session cookie is required; no JWT requirement.public: no end-user JWT/session requirement.
Session behavior is governed by gateway/security configuration (same-site policy, TTL, and issued-at skew tolerance).
Routes using dual or cookie-only automatically enable CSRF protection using the double submit cookie pattern. The gateway issues a non-HttpOnly X-CSRF-Token cookie and the frontend client must echo that value back in the X-CSRF-Token request header.
Note: When the session cookie has been set to
SameSite: None, the CSRF token is also sent on a response header,X-CSRF-Token, because it is assumed the developer wants cross-site functionality. In that case, the token value should be read from the header rather than attempting to read the cookie.
Example client setup:
const token = Cookie.get('X-CSRF-Token');
const instance = axios.create({
headers: {'X-CSRF-Token': token}
});CSRF behavior for dual and cookie-only routes:
GET,HEAD, andOPTIONSare treated as safe methods and are not subject to CSRF validation.- Other methods are unsafe and must include a matching
X-CSRF-Tokencookie/header pair. - A new CSRF cookie is returned for each request.
- Each CSRF cookie expires after
7200seconds. - A safe
GETcan be used to obtain a fresh CSRF cookie before making an unsafe request such asPOST.
Client API Keys
Native Z-Arch clients are declared in clients and mapped to API-key identities stored in secret manager.
Each client includes:
idapi_keysecret referenceadd_to_edgeto control edge-side API-key injection on the client segment
Firebase app provisioning is separate and declared under authn.firebase.apps.
Capability Enforcement Status
Z-Arch includes authz capability and role structures with route-level requires checks. This model exists in the contract surface but remains an actively developing area.
Perimeter vs Domain Authorization
Perimeter authentication and route policy are enforced by the edge/gateway boundary. Domain authorization remains a service concern.
Serverless Primitives
Z-Arch breaks backend architecture into explicit primitives.
Services
Services are request/response workloads that can be exposed as API endpoints or kept internal.
Common use cases:
- Product API endpoints.
- Session and account services.
- Internal orchestration APIs.
Example:
services:
- id: orders
endpoint: true
routes:
- path: /orders
method: GET
auth: dual
- path: /orders
method: POST
auth: dual
location: /services/orders
env:
LOG_LEVEL: info
secrets:
- id: database-password
mount: DATABASE_PASSWORD
- id: orders-tls-certificate
version: 3
mount: /var/run/secrets/orders/tls.crt
flags: []
targets: [payments, order-events]Jobs
Jobs are run-to-completion workloads for asynchronous or batch execution.
Common use cases:
- Daily reconciliation.
- Data cleanup.
- Bulk import/export.
Example:
jobs:
- id: daily-reconcile
location: /jobs/daily-reconcile
env:
BATCH_SIZE: "500"
secrets:
- id: reporting-token
version: production
mount: REPORTING_TOKEN
flags:
- "--task-timeout=1200s"
targets: [orders, reporting-events]For source deployments, location is resolved relative to zarch.yaml. The service or job directory must contain a file named Dockerfile; the CLI stops the deployment with an error when it is missing. Use image instead of location to deploy a prebuilt container image.
Runtime Secrets
Services and jobs can declare Secret Manager references with the optional secrets list. Each item requires an id matching the secret name and a mount. A mount beginning with / is an absolute file path; any other mount is the runtime environment variable name. version may be an enabled positive version number or Secret Manager alias and defaults to latest when omitted. Aliases must start with a letter, may contain letters, numbers, hyphens, and underscores, and may be at most 63 characters long.
services:
- id: orders
secrets:
- id: database-password
mount: DATABASE_PASSWORD
- id: tls-certificate
version: 3
mount: /var/run/secrets/tls.crt
jobs:
- id: daily-reconcile
secrets:
- id: reporting-token
version: production
mount: REPORTING_TOKENAt deployment, Z-Arch verifies that every configured version resolves to an enabled Secret Manager version, grants the resource’s runtime service account roles/secretmanager.secretAccessor on each referenced secret, and supplies the references to Cloud Run through its managed --update-secrets option. Secret-related deployment flags are not accepted in flags.
In interactive mode, when an omitted or latest version is unavailable, the CLI securely prompts for the secret value and creates a new enabled version. The value is masked and is never written to zarch.yaml, command arguments, review output, lifecycle payloads, or logs. An unavailable explicit version or alias always fails. Non-interactive deployments fail for every unavailable version instead of prompting.
--update-secrets is incremental. Removing an entry from zarch.yaml does not remove a secret binding already present on Cloud Run; remove stale runtime bindings explicitly before or alongside the configuration change.
Topics
Topics are passive event channels for publish/subscribe flows.
Common use cases:
- Event fan-out.
- Workflow decoupling.
- Retry-friendly async processing.
Example:
topics:
- id: reporting-events
sub:
- id: analytics
mode: push
ack_deadline_secs: 60Schedulers
Schedulers are time-based trigger resources that invoke services, run jobs, or publish messages.
Common use cases:
- Hourly sync tasks.
- Nightly cleanup.
- Weekly report generation.
Example:
schedulers:
- id: hourly-sync
cadence:
every: 1
unit: hour
timezone: "UTC"
targets: [daily-reconcile]Targets and Trust Edges
targets define the declared execution graph for a project.
- Service and job targets define callable relationships.
- Topic targets define event publish permissions.
- Scheduler targets define trigger actions.
Trust edges are explicit and deterministic:
- Short-lived assertions are minted only for declared call paths.
- Receivers validate caller identity and intended audience.
- Undeclared caller/target edges are outside the trust graph and are not accepted as valid internal calls.
This keeps topology reviewable and prevents hidden trust relationships from emerging outside zarch.yaml.
II. Operating Guides
Getting Started
Project Bootstrap
Create a new app:
zarch new appTypical guided flow:
- Choose new vs existing repository mode.
- Select template/source repository.
- Select cloud project and region details.
- Configure domain and edge deployment options.
- Configure authentication options.
- Apply bootstrap and initial deployment steps.
Existing Repository Bootstrap
You can bootstrap into an existing repository/branch instead of generating from a template repository.
zarch new app --not-new-repoUse this when:
- You already have an active codebase integrated with Z-Arch.
- You want to bootstrap a separate GCP project for different branches (e.g. prod, test)
- You are integrating an existing codebase with Z-Arch for the first time.
New Project from Any Template
You can point Z-Arch to any GitHub template repository that includes a valid zarch.yaml.
zarch new app --template-repo owner/template-repoThis enables:
- Company-standard starter templates.
- Shared community templates.
- Distributing / installing ready to deploy applications.
Sharing Templates
Recommended template practice:
- Keep a stable template repository with a clean
zarch.yamlbaseline. - Version template changes via standard branch/tag workflows.
- Document template assumptions in repository README.
- Bootstrap new projects by referencing that template repository directly.
GitHub Prerequisites
Z-Arch installs a project-specific, read-only SSH deploy key on the connected GitHub repository so the Control Plane and Extension Runner can clone it without receiving your human GitHub credential. The private half of this key is stored in the project’s Secret Manager.
Confirm that the organization or enterprise that owns the repository permits deploy keys. If deploy keys are disabled, Z-Arch cannot install the repository key and project bootstrap cannot complete. Organization owners can review this policy under Organization Settings → Deploy keys; an enterprise policy may prevent organization owners from enabling it.
Deployment Model
Z-Arch operates on a convergent deployment model. You declare intent in zarch.yaml; deployment applies that intent in a fixed order; repeated deployment converges toward the same runtime shape.
Do not manually orchestrate infrastructure components. The Control Plane applies resources in dependency order.
Deployment Model
When deploying a full project, resources are applied in this order:
- Topics
- Services
- Jobs
- Schedulers
- Gateway
- Edge updates
This ordering ensures:
- Event infrastructure exists before publishers or subscribers.
- Services and jobs exist before schedulers reference them.
- The Gateway reflects the current route definitions.
- Edge routing reflects the current Gateway endpoints.
Most workflows use:
zarch deploy allThis applies the full convergence cycle.
Anti-Pattern
- Manually wire load balancers.
- Manually configure IAM bindings for service trust.
- Manually synchronize route definitions between services and ingress.
- Deploy hidden infrastructure outside
zarch.yaml.
Operational changes are expressed in configuration. Deployment enforces alignment.
Convergence Principle
Repeated deployment of the same configuration produces the same infrastructure state.
Drift introduced outside Z-Arch is not part of the intended model. The Control Plane is responsible for re-aligning declared resources with runtime state.
Devbox
Overview
zarch devbox gives each developer a project-scoped cloud workspace that can be created, paused, resumed, and replaced on demand.
Devboxes use Identity-Aware Proxy (IAP) by default: SSH is tunneled through Google Cloud and the VM has no external IP address.
For teams, devboxes provide:
- Faster onboarding with fewer machine-specific setup issues.
- Lower local environment drift across engineers.
- Better isolation between projects and developers.
- Safer experimentation with limited blast radius.
- Cost control through explicit start/stop lifecycle commands.
Why Ephemeral Environments Matter
Devboxes are intentionally disposable. If an environment becomes unstable, slow, or misconfigured, the preferred recovery path is to replace it quickly instead of spending hours on manual repair.
This model improves:
- Mean time to recovery for developer blockers.
- Reproducibility across engineers.
- Support handoff quality (shared, repeatable recovery workflow).
Quick Start
# 1) Set active project directory
zarch set project /path/to/your/zarch/project
# 2) Authenticate the local gcloud CLI used for provisioning and SSH tunneling
gcloud auth login
gcloud config set project <project-id>
# 3) Create an IAP devbox (the default)
zarch new devbox <username>
# 4) Connect over SSH using the alias printed by the CLI
ssh devbox-<project-id>-<usr3>Example:
zarch new devbox ram
ssh devbox-myproject-ramPrerequisites
Before creating or managing IAP devboxes, ensure:
- You are in a valid Z-Arch project directory (
zarch.yamlexists). - Your active GCP project is configured and accessible.
- The computer running
zarchand SSH has an authenticated localgcloudidentity. Cloudflare credentials and a project domain are not required for the IAP default.
Required IAM capabilities typically include:
- Compute Engine instance management
- IAM service account and policy binding management
- IAM service-account and IAP policy binding management
- Service API enablement permissions
Access Modes and IAP Principals
zarch new devbox [username] --access-mode [iap|cloudflare] --iap-principal PRINCIPAL--access-mode is case-insensitive and defaults to iap. For IAP, a plain --iap-principal email becomes user:<email>; explicit user:, group:, and serviceAccount: principals are accepted. If omitted, Z-Arch uses the active account from local gcloud auth list. If no active account exists, provide --iap-principal.
--iap-principal is invalid with --access-mode cloudflare.
IAP authorization and SSH authentication
IAP authorization and SSH authentication are separate controls. The selected principal may open that instance’s IAP tunnel only on port 22 and receives Z-Arch’s minimal project instance-read role; it does not receive project-wide IAP access. The Z-Arch-generated ED25519 key then authenticates the custom Linux user. enable-oslogin=FALSE remains set, so tunnel authorization never replaces the SSH key and a valid key cannot bypass IAP authorization.
Networking and Shared Resources
Devboxes use the shared custom-mode VPC zarch-devbox-net, separate from the default VPC and application zarch-net. Each used region gets a managed /20 subnet from 10.64.0.0/12.
- IAP mode: internal IPv4 only, a tagged IAP SSH firewall rule, and shared regional Cloud NAT for outbound downloads.
- Cloudflare mode: Premium ephemeral external IPv4, a tagged public SSH firewall rule, and Cloudflare DNS forwarding.
Managed names are zarch-devbox-subnet-<region> for the subnet; zarch-devbox-allow-iap-ssh with tag zarch-devbox-iap for IAP; zarch-devbox-allow-public-ssh with tag zarch-devbox-cloudflare for Cloudflare; and, for IAP egress, zarch-devbox-router-<region> and zarch-devbox-nat-<region>.
The DEVBOX VPC, regional subnets, firewalls, routers, NATs, and dedicated IAP instance-read role are project-shared resources. Deleting an individual devbox never deletes them.
Optional Cloudflare Mode
Cloudflare is an explicit public SSH/DNS mode. It requires a configured project domain, a Cloudflare token, Secret Manager access, and DNS-update permission:
zarch connect cloudflare
zarch new devbox alice --access-mode cloudflareFirst Login: Identity and Authorization (Recommended Standard)
After your first SSH login, it is highly recommended that you authenticate gcloud and ADC as your user identity.
Why this matters
The VM service account is intentionally limited for platform runtime operations.
Developer workflows (project administration, many gcloud actions, and parts of Z-Arch usage) should run under the developer’s user identity for correct authorization and auditability.
Run this once per new devbox
gcloud auth login
gcloud auth application-default login
gcloud config set project <project_id>Verify
gcloud auth list
gcloud auth application-default print-access-tokenExpected outcome:
- Your interactive CLI/API calls execute as your user identity.
- You avoid common
PERMISSION_DENIEDfailures caused by relying on the VM service account for user workflows.
Default Tooling Baseline
New devboxes are delivered with a platform baseline of developer tooling.
This baseline combines platform-provisioned tools and Ubuntu image utilities.
Platform baseline tools
The following tools are expected in a newly provisioned devbox:
zarch(Z-Arch CLI)gcloud(Google Cloud CLI)gitgh(GitHub CLI)curljqbinutilsbuild-essentialdockerand Docker Compose plugin- Node.js (LTS) and
npm wranglerfirebaseCLI- Go (
go, installed viagolang-go) - Rust (
rustcandcargo) codexCLIclaudeCLIgeminiCLIplaywrightandplaywright-mcp- ‘python3-venv and python-is-python3’ Python environment tools
Ubuntu base image utilities
In addition to platform tooling, Ubuntu provides core operational utilities commonly used for day-to-day engineering tasks, including:
- shell and core GNU tooling (
bash, coreutils) - package management (
apt) - native build and linker tooling (
binutils,build-essential) - language toolchains installed from Ubuntu packages (
golang-go,rustc,cargo) - service management (
systemctl/systemd) - SSL and crypto utilities (
ca-certificates,gnupg) - common Linux networking and process utilities
Validation checklist (recommended)
After first login, validate your baseline:
command -v zarch gcloud git gh curl jq docker node npm wrangler go rustc cargoIf a required tool is missing, follow your team’s custom provisioning approach in the next section.
Command Reference
| Command | Purpose |
|---|---|
zarch new devbox [username] [--access-mode iap|cloudflare] [--iap-principal PRINCIPAL] | Create a devbox; IAP is the default. |
zarch devbox list | List devboxes with ACCESS MODE and INTERNAL IP; Cloudflare rows also retain NAT IP and forward domain. |
zarch devbox on <username> | Start a devbox VM by username. |
zarch devbox off <username> | Stop a devbox VM by username. |
zarch devbox delete <username> | Delete a devbox VM without deleting shared DEVBOX networking or the shared IAP read role. |
zarch devbox delete <username> --delete-service-account | Delete VM and associated service account. |
Custom Provisioning with an Appended Startup Script
If your team needs additional packages, internal CLIs, or project-specific setup, you can provide an additional startup script during devbox creation.
How it works
During zarch new devbox <username> (interactive mode), Z-Arch prompts for:
Additional startup script path (optional)
Provide a local script file path. Z-Arch will append and run that script as part of devbox provisioning.
Recommended usage pattern
- Store team customization scripts in source control (for example:
devbox/custom-startup.sh). - Keep scripts idempotent so reruns are safe.
- Use explicit version pins for critical toolchains.
- Log clearly to simplify support and audits.
Example flow
zarch new devbox alice
# When prompted:
# Additional startup script path (optional): ./devbox/custom-startup.shImportant notes
- The script is applied at provisioning time for that devbox creation.
- To apply a changed script to an existing environment, the recommended pattern is recreate:
zarch devbox off <username>zarch devbox delete <username>zarch new devbox <username>(with updated script path)
Naming and Access Conventions
For username alice in project my-project:
- VM name:
devbox-alice - Service account:
devbox-alice-sa@my-project.iam.gserviceaccount.com - Local SSH alias format:
devbox-<project-id>-<usr3>
Only Cloudflare devboxes have a forward domain such as alice.dev.<domain>. IAP aliases use a local SSH ProxyCommand with gcloud compute start-iap-tunnel; no --account is pinned, so the active local gcloud identity is used.
Z-Arch updates local SSH configuration on the machine running zarch, enabling alias-based connection. The same alias works with VS Code Remote SSH, scp, sftp, and rsync.
Ephemeral Operations Model
Standard policy
- One devbox per developer per project.
- Prefer fast replacement over prolonged repair when the environment is degraded.
- Treat devbox state as recoverable and reproducible.
Reset runbook (recommended for broken environments)
zarch devbox off <username>
zarch devbox delete <username>
zarch new devbox <username>Then reconnect:
ssh devbox-<project-id>-<usr3>Operational Runbooks
Daily startup / shutdown
Start your environment:
zarch devbox on <username>Stop when idle:
zarch devbox off <username>Environment recovery (nuke and recreate)
Use this when troubleshooting exceeds a reasonable threshold:
zarch devbox off <username>
zarch devbox delete <username>
zarch new devbox <username>Developer offboarding
Remove compute resource:
zarch devbox delete <username>If identity should also be removed:
zarch devbox delete <username> --delete-service-accountIncident triage flow
- Verify VM state:
zarch devbox list- Verify reachability:
- SSH alias in
~/.ssh/config - IAP: local
gcloud auth listshows the intended identity - Cloudflare: forward domain resolves as expected
- If unresolved after basic checks, escalate to reset runbook.
Team Operating Model
Recommended team practices:
- Use stable, unique usernames per developer.
- Define ownership boundaries: one devbox equals one responsible engineer.
- Use start/stop discipline for cost management.
- Standardize recovery on replace-over-repair for severe drift.
- Use project-scoped credentials wherever possible for governance clarity.
Security and Identity Notes
- Devboxes use least-privilege runtime identity for platform-controlled operations.
- Developer-admin actions should use authenticated user identity (see ADC section).
- Runtime secret access is managed through project controls, not local plaintext artifacts.
- Devbox lifecycle actions do not require adding new fields to
zarch.yaml. - Start and stop are transport-neutral. On IAP deletion, the principal’s shared read-role binding is removed only when no remaining IAP devbox references it; shared VPC, subnet, firewall, router, NAT, and custom-role resources remain intact.
Troubleshooting
| Symptom | Likely Cause | Resolution | Escalation |
|---|---|---|---|
No valid zarch.yaml found in this project directory. | Wrong working directory or project not set | zarch set project /path/to/project | Confirm repo contains correct zarch.yaml |
| IAP create says no active identity, or the wrong account is used | Local gcloud login expired or a different account is active | Run gcloud auth login, verify gcloud auth list, or use --iap-principal | Confirm the intended normalized principal |
| IAP tunnel is denied | Selected principal lacks instance IAP permission or its minimal lookup role | Review the instance-specific IAP policy and Z-Arch DEVBOX read-role binding | Ask a project administrator to correct the binding |
| IAP tunnel cannot reach port 22 | IAP firewall/tag coverage is missing or shared resources drifted | Check instance tags and the managed IAP firewall | Correct incompatible shared resources before retrying |
| Outbound package downloads fail in IAP mode | Regional NAT/router coverage is missing or incompatible | Verify Cloud NAT covers the DEVBOX subnet primary range | Review the managed router/NAT configuration |
SSH ProxyCommand fails | Local gcloud is unavailable, unauthenticated, or cannot start the tunnel | Verify gcloud is on PATH, run gcloud auth list, then retry ssh -vvv <alias> | Review local IAP permissions and tunnel output |
| SSH key is rejected after the IAP tunnel opens | Wrong alias/key, changed key permissions, or custom-user mismatch | Verify IdentityFile, Linux user, and key permissions; keep OS Login disabled | Recreate after correcting key material if needed |
| Cloudflare token or DNS update fails | Cloudflare mode lacks a domain, token, Secret Manager access, or DNS permission | Run zarch connect cloudflare and verify the project domain/token scope | Retry explicitly with --access-mode cloudflare; IAP devboxes are unaffected |
PERMISSION_DENIED running gcloud/Z-Arch in VM | Running as limited VM identity instead of user | Run first-login identity steps (gcloud auth login, gcloud auth application-default login, set project) | Verify org/project IAM grants for user |
Authenticated as VM service account when user actions are expected | User auth/ADC not initialized | Re-run user auth + ADC commands | Recheck gcloud auth list output |
| SSH alias fails to connect | VM is off or alias is missing; IAP may have a local tunnel/identity failure, while Cloudflare may have a DNS-forwarding failure | zarch devbox list, zarch devbox on <username>, verify ~/.ssh/config; check IAP gcloud auth list or the Cloudflare forward domain for the selected mode | If unresolved, run reset workflow |
| Devbox repeatedly unstable after manual fixes | Environment drift or corrupted local state | Execute reset runbook (delete + recreate) | Escalate with logs and issue context |
Success Criteria
Your devbox adoption is working well when:
- New developers can reach a usable environment quickly.
- Most severe environment failures are resolved by fast recreate cycles.
- Daily operations use
on/offpredictably for cost control. - Teams see reduced “works on my machine” inconsistency.
- The local IAP identity and SSH key are consistently configured for each devbox.
Day-2 Operations
Incremental Changes
Z-Arch supports targeted deployments. You do not redeploy the entire system for every change.
Changing a Service
If you modify:
- Service code
- Environment variables
- Runtime secret references
- Declared targets
You redeploy that service:
zarch deploy service <name>If you modify routes for that service, the Gateway must also be redeployed so route definitions remain aligned.
Changing Gateway Behavior
If you modify:
- Route auth modes
- Session configuration
- CORS configuration
- Security settings
You redeploy the Gateway:
zarch deploy gatewayIf rotating session encryption keys:
zarch deploy gateway --rotate-session-keySession rotation invalidates existing sessions by design.
Client Identity Changes
Client API keys can be regenerated independently:
zarch client all
zarch deploy edgeRemoving an API key’s name from the api_key field in a client’s config before running zarch client all will rotate the secret.
Troubleshooting Guide
Common issues and checks:
- Deploy keys are disabled.
- In GitHub, check Organization Settings → Deploy keys. An enterprise policy may prevent organization owners from enabling deploy keys. See GitHub Prerequisites.
- Active project not set.
- Run
zarch set project <path>.
- Run
- Cloud project not active.
- Verify cloud CLI auth and active project selection.
- Deploy command reports missing resource ID.
- Confirm
idexists in the relevant block.
- Confirm
- Route validation errors.
- Verify path format and duplicate method/path keys.
- Runtime secret validation errors.
- Confirm every
id:versionexists and is enabled, mounts are unique, and secret options are declared undersecretsrather thanflags.
- Confirm every
- CORS failures in browser clients.
- Verify
security.cors.allowed_origins, headers, methods, and credential settings.
- Verify
- Auth failures on protected endpoints.
- Check route
authmode and your JWT/session flow assumptions.
- Check route
Multi-Region Model
Use regions in zarch.yaml to declare multi-region intent, then apply region-aware edge load-balancing settings.
Regions are intended to be convergent. Divergent per-region code paths and undeclared regional drift are outside the intended model.
In practice:
- Deploy intent remains centralized in one
zarch.yamlcontract. - Gateway and service shape are expected to remain structurally consistent across declared regions.
- Edge load-balancing policy determines request distribution across those converged regional deployments.
Security Posture and Shared Responsibility
Z-Arch guarantees perimeter authentication enforcement at the edge/gateway boundary according to declared route policy.
Z-Arch does not guarantee domain authorization correctness inside your application services.
Z-Arch does not protect against incorrect business logic, unsafe data handling, or domain-level policy mistakes in service code.
Security best-practice baseline:
- Keep
zarch.yamldeclarative and free of sensitive values. - Prefer explicit origin allowlists for CORS.
- Prefer
dualfor browser-authenticated routes unless there is a deliberate reason to usejwt-onlyorcookie-only. - When using
dualorcookie-only, ensure clients echo theX-CSRF-Tokencookie into theX-CSRF-Tokenrequest header and allow that header through CORS. - Treat API keys as perimeter credentials and rotate them periodically.
- Keep service
targetsminimal and intentional. - Use short-lived internal trust tokens and avoid custom long-lived shared secrets between services.
- Always keep secrets in Secret Manager.
- Call third-party APIs from backend services, not directly from untrusted clients where avoidable.
III. Reference Manuals
CLI Reference
Global Usage
zarch [--non-interactive] [--quiet] <command>Run with no command to open the interactive console:
zarchBehavior:
zarchwith no subcommand opens the interactive shell.--non-interactiveavoids prompts and relies on explicit flags and existing config values.--quietsuppresses most non-critical output.
Command groups:
| Group | Purpose |
|---|---|
set | Active project and Control Plane source pinning |
new | App bootstrap and Devbox creation |
deploy | Resource deployments |
devbox | Devbox lifecycle operations |
client | API client identity and key generation |
connect / disconnect | Third-party credential management |
ext | Extension scaffolding, install, and hook triggering |
docs | Read the Z-Arch documentation table of contents or a named section in the terminal |
mcp | Local MCP server for controlled config operations and AI assisted development |
update | Update the Control Plane to the latest version |
register | Register your project with a lifetime Z-Arch license |
set
Project and Control Plane pinning commands.
zarch set project [path]
zarch set branch <branch>
zarch set repo <owner/repo>
zarch set defaultsUsage notes:
set projectactivates a local directory that containszarch.yaml.set branchchanges the branch pinned for Control Plane operations.set repochanges the repository pinned for Control Plane operations.set defaultsis reserved for default preference management.
new
Create new platform resources.
zarch new app [options] [path]
zarch new devbox [username] [--access-mode iap|cloudflare] [--iap-principal PRINCIPAL]new app options:
| Option | Purpose |
|---|---|
--region | Initial deployment region |
--project-id-prefix | Prefix for generated cloud project IDs |
--template-repo | Source template repository (owner/repo) |
--new-repo | New repository name |
--private-repo | Create repository as private |
--signup-mode | Signup policy |
--auth-methods | Managed auth methods |
--mfa | MFA mode |
--billing-account | Billing account override |
--not-new-repo | Use an existing repository |
Examples:
zarch new app --template-repo my-org/my-template
zarch new app --region us-east1 --project-id-prefix prod-
zarch new app --not-new-repo
zarch new devbox alice
zarch new devbox alice --access-mode cloudflaredeploy
Deploy one or more declared resources.
zarch deploy gateway [--region] [--rotate-session-key]
zarch deploy service [--region] <name>
zarch deploy job [--region] <name>
zarch deploy topic [--region] <name>
zarch deploy scheduler [--region] <name>
zarch deploy edge
zarch deploy all [--region]Usage notes:
- If
--regionis omitted, deploy commands iterate through all regions declared in config. deploy allapplies the full deployment sequence for declared resources.--rotate-session-keyon gateway deployment rotates session encryption key material and forces new sessions.
Examples:
zarch deploy service orders
zarch deploy gateway --region us-east1 --rotate-session-key
zarch deploy alldevbox
Manage Devbox lifecycle.
New devboxes default to IAP. list reports access mode and internal IP; Cloudflare devboxes additionally report their NAT IP and forward domain. Deleting an instance does not delete shared DEVBOX networking, NAT, firewall, router, or custom-role resources.
zarch devbox list
zarch devbox on <username>
zarch devbox off <username>
zarch devbox delete <username> [--delete-service-account]client
Manage API clients and keys.
zarch client new <id> [--no-edge]
zarch client all
zarch firebase app new <platform> <client>
zarch firebase app allUsage notes:
- Native clients are gateway/API-key identities and do not have a platform
type. --no-edgeprevents automatic key injection setup at the Edge Proxy layer.- Firebase app provisioning is explicit and uses
authn.firebase.apps.
Examples:
zarch client new web
zarch firebase app new web web
zarch client allconnect and disconnect
Manage third-party credentials.
zarch connect github
zarch connect cloudflare [--project|--global]
zarch disconnect github
zarch disconnect cloudflare [--project|--global]Scope behavior:
- GitHub is one local account-level encrypted-vault credential,
github; it has no project or global scope and is not synchronized tozarch_secrets. - Cloudflare defaults to the active project.
--projectis an explicit equivalent. - Cloudflare
--globalcreates or removes a legacy global fallback token without requiring an active project. This remains supported for compatibility but is deprecated for normal workflows. - When an interactive
zarch new appenables the edge proxy and no exact project-scoped Cloudflare token exists, Z-Arch opens a prefilled account-token creation form automatically. A legacy global token does not suppress this onboarding. - Select the Cloudflare account that owns the project resources and restrict Zone Resources to the prompted root domain. Creating account-owned tokens requires suitable Cloudflare account administration privileges. See Cloudflare account API tokens and token template URLs.
- Non-interactive project bootstrap cannot prompt for a token. Connect the project token with
zarch connect cloudflare, then runzarch deploy edge.
ext
Manage Z-Arch Extensions.
zarch ext new <name>
zarch ext install <source> [--all]
zarch ext trigger <hook_name> [<extension_block>...] [--region <region>]Usage notes:
ext newscaffolds a new extension package.ext install --allinstalls all discoverable extensions in the active project.ext triggermanually dispatches a lifecycle hook for testing and operations.
mcp
Start the local MCP server.
zarch mcpExposed MCP tools are focused on safe, validated config operations, scaffolding workflows, and authoritative Z-Arch documentation. You do not need to manually invoke this command.
docs
Read the Z-Arch documentation directly in the terminal.
zarch docs
zarch docs '<SECTION>'Usage notes:
- Run
zarch docsfirst to display the table of contents and discover section names. - Pass a section heading to render that section. Matching is case-insensitive.
- Quote section names that contain spaces, for example
zarch docs 'Platform Architecture'. - Documentation is retrieved remotely, so this command requires network access.
update
Update the deployed Control Plane environment.
zarch updateregister
Purchase a Z-Arch license. Scope = license type.
zarch register [scope]zarch.yaml Reference
zarch.yaml is the source of truth for your platform architecture:
- It defines what exists.
- It defines what is exposed.
- It defines how resources are allowed to interact.
Compressed Architectural Language
zarch.yaml is not just configuration. It is a compact mapping of your backend architecture.
In one machine-legible document, it encodes:
- Identity and authentication boundaries.
- Authorization capabilities and roles.
- IAM-relevant topology via explicit resource targets.
- Infrastructure primitives (Services, Jobs, Topics, Schedulers).
- Gateway exposure and route behavior.
- Extension declarations and deployment controls.
- Operational guardrails and runtime policy defaults.
Core architectural intent is centralized and structured. This compactness is useful for both engineers and LLMs:
- Humans can review architecture intent quickly without chasing hidden dashboard state.
- The grammar is small enough for LLMs to reason over the full backend topology without burning tokens or hallucinating.
Z-Arch is therefore opinionated, structured, and machine-reasonable by design: a deterministic serverless pattern expressed in a concise file that is practical for both production operations and AI-assisted engineering.
Top-Level Contract
Top-level keys:
| Key | Required | Type | Notes |
|---|---|---|---|
platform | Yes | string | Currently gcp |
domain | Yes | string | null |
clients | Yes | array | Native gateway/API-key clients |
gateway | Yes | object | null |
authn | Yes | object | null |
security | Yes | object | null |
project_id | No | string | null |
regions | No | array[string] | Deployment regions |
edge | No | object | null |
authz | No | object | null |
services | No | array | Service resources |
jobs | No | array | null |
topics | No | array | null |
schedulers | No | array | null |
Gateway and Security Blocks
Gateway block:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
gateway.type | string | null |
gateway.version | string | null |
gateway.min_instance | integer | 0+ |
gateway.session.service_id | string | null |
gateway.session.stateful | boolean | Stateful session toggle |
gateway.session.login_iat_skew_secs | integer | Session login JWT issued-at skew tolerance, 0+ |
gateway.session.cookie.samesite | string | Strict, Lax, None |
gateway.session.cookie.ttl_secs | integer | Session cookie lifetime seconds, 0+ |
gateway.session.cookie.domain | string | Optional cookie domain override. Defaults to domain when omitted. |
Security block:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
security.cors.allowed_origins | array[string] | Explicit allowed origins |
security.cors.allowed_headers | array[string] | Allowed request headers. Include X-CSRF-Token for browser clients using dual or cookie-only routes. |
security.cors.allowed_methods | array[string] | GET, POST, PUT, PATCH, DELETE, OPTIONS, HEAD, TRACE, CONNECT |
security.cors.expose_headers | array[string] | Headers exposed to browser JS |
security.cors.allow_credentials | boolean | Credentialed browser requests |
security.cors.max_age_seconds | integer | Preflight cache duration |
Auth Blocks
authn supports two primary models:
- Managed Firebase bootstrap path (
authn.firebase.auth_methods,authn.firebase.signup_mode,authn.firebase.mfa). - Generic OIDC validation path (
discovery,client_id).
authn field values:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
authn.discovery | string | null |
authn.client_id | string | null |
authn.firebase.project_id | string | null |
authn.firebase.tenant_id | string | null |
authn.firebase.auth_methods | array[string] | Email/Password, Email Link, Google, Github, Microsoft, Apple, Facebook |
authn.firebase.signup_mode | string | null |
authn.firebase.mfa | string | null |
authn.firebase.apps[].client | string | Existing clients[].id that receives the Firebase config packet |
authn.firebase.apps[].platform | string | web, ios, android |
authz supports:
capabilitiesrolesroles.<role>.grants
Use requires on routes for capability checks where applicable.
Resource Blocks
clients:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
id | string | Unique client ID |
api_key | string | null |
add_to_edge | boolean | Edge-side key injection toggle |
services:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
id | string | Unique service ID; maximum 27 characters |
endpoint | boolean | Expose via Gateway |
routes | array | Route definitions |
location | string | null |
image | string | null |
env | object | Runtime env vars |
flags | array[string] | Additional deploy flags |
targets | array[string] | Allowed downstream targets |
timeout.connect | integer | Service connection timeout in seconds for all routes in the service. Default 60 |
timeout.send | integer | Request write timeout in seconds for all routes in the service. Default 60 |
timeout.read | integer | Response read timeout in seconds for all routes in the service. Default 60 |
location and image are mutually exclusive. A service using location is a source deployment and must have a Dockerfile in the resolved directory.
Service route fields:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
path | string | Literal (/path) or regex (^/path$) |
method | string | HTTP verb |
mode | string | http, stream, webhook, websocket |
auth | string | dual, jwt-only, cookie-only, public |
content_type | string | Required request content type |
rate | integer | Route limit, minimum 1 |
requires | array[string] | Capability checks |
timeout.connect | integer | Service connection timeout in seconds. Takes precedence over service level timeout |
timeout.send | integer | Request write timeout in seconds. Takes precedence over service level timeout |
timeout.read | integer | Response read timeout in seconds. Takes precedence over service level timeout |
health | boolean | Health endpoint designation |
Every route must define auth explicitly. When auth is dual or cookie-only, the route automatically uses CSRF protection through the double submit cookie method. Browser clients must send the X-CSRF-Token cookie value back in an X-CSRF-Token request header.
jobs:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
id | string | Unique job ID; maximum 27 characters |
location | string | null |
image | string | null |
env | object | Runtime env vars |
flags | array[string] | Additional deploy flags |
targets | array[string] | Callable/publish targets |
A job using location is a source deployment and must have a Dockerfile in the resolved directory. Use image for a prebuilt job image.
topics:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
id | string | Unique topic ID |
sub | array | null |
Subscriber entry forms:
- String form: subscriber ID.
- Object form:
id, optionalmode(push/pull), optionalack_deadline_secs, optionalmax_delivery_attempts, optionaldead_letter_topic.
schedulers:
| Field | Type | Allowed Values / Behavior |
|---|---|---|
id | string | Unique scheduler ID; maximum 27 characters |
cadence | object | null |
cron | string | null |
paused | boolean | Pause state |
targets | array[string] | Trigger targets |
Scheduler rule:
- Exactly one of
cadenceorcronmust be set.
Cadence fields:
everyunit(minute,hour,day,week,month)- optional
at - optional
on - optional
timezone
Full Example
platform: gcp
project_id: my-project
domain: api.example.com
regions:
- us-east1
edge:
load_balancer:
method: DIRECT
rate_limiter: null
clients:
- id: web
api_key: API_KEY_WEB
add_to_edge: true
gateway:
type: serverless
min_instance: 0
session:
service_id: session
stateful: false
login_iat_skew_secs: 60
cookie:
samesite: Lax
ttl_secs: 604800
# domain: app.example.com
authn:
discovery: https://issuer.example.com/.well-known/openid-configuration
client_id: my-client-id
firebase:
project_id: my-project
tenant_id: null
auth_methods: []
signup_mode: null
mfa: null
apps:
- client: web
platform: web
authz:
capabilities:
- orders:read
- orders:write
roles:
admin:
grants:
- orders:read
- orders:write
security:
cors:
allowed_origins:
- https://app.example.com
allowed_headers:
- Authorization
- Content-Type
- X-CSRF-Token
- x-api-key
allowed_methods:
- GET
- POST
- PUT
- PATCH
- DELETE
expose_headers: []
allow_credentials: true
max_age_seconds: 600
services:
- id: session
endpoint: true
routes:
- path: /session
method: GET
auth: dual
- path: /session/login
method: POST
auth: jwt-only
- path: /session/logout
method: POST
auth: dual
- path: /session/health
method: GET
auth: public
health: true
location: /services/session
env: {}
secrets:
- id: session-key
mount: /zarch/mounts/session_key/session_key
flags: []
targets: []
- id: orders
endpoint: true
routes:
- path: /orders
method: GET
auth: dual
requires: [orders:read]
- path: /orders
method: POST
auth: dual
requires: [orders:write]
- path: ^/orders/[a-zA-Z0-9_-]+$
method: GET
auth: dual
location: /services/orders
env: {}
secrets:
- id: database-password
mount: DATABASE_PASSWORD
flags: []
targets: [order-events]
jobs:
- id: nightly-sync
location: /jobs/nightly-sync
env: {}
secrets:
- id: reporting-token
version: production
mount: REPORTING_TOKEN
flags:
- "--task-timeout=1200s"
targets: [orders]
topics:
- id: order-events
sub:
- id: orders
mode: push
schedulers:
- id: nightly-sync-scheduler
cron: "0 2 * * *"
paused: false
targets: [nightly-sync]Route Semantics
Path Semantics
Path rules:
- Literal paths must start with
/and contain no whitespace. - Regex paths are opt-in and must start with
^/and end with$.
Method rules:
- Methods must be valid HTTP verbs supported by the config contract.
Uniqueness rules:
- Duplicate route keys are rejected within the same resource for the same method/path kind.
- Resource IDs must be unique across major resource groups.
- Service, job, and scheduler IDs must be 27 characters or fewer.
authz.roles.*.grantsvalues must exist inauthz.capabilities.
Auth Modes
Route auth mode is declared per route:
dual- JWT and session cookie required.jwt-only- JWT required, no session cookie.cookie-only- session cookie required, no JWT.public- no end-user authentication required.
Routes using dual or cookie-only automatically enable CSRF protection through the double submit cookie method. The gateway returns a non-HttpOnly X-CSRF-Token cookie and expects the same value in the X-CSRF-Token request header. GET, HEAD, and OPTIONS are safe methods and are not subject to CSRF validation. Other methods are unsafe, receive a fresh CSRF cookie on each request, and must present a matching cookie/header pair. Each CSRF cookie expires after 7200 seconds.
public routes remain behind gateway perimeter controls and are not direct service ingress endpoints.
Capability Requirements
Where authz is configured, route requires declarations enforce capability checks as part of gateway policy evaluation.
Conditional Requirements (required_when)
Route fields follow conditional requirement semantics in the config contract:
pathandmethodare required for each route entry.auth,content_type,rate,requires, andhealthare conditional policy modifiers.- Conditional requirement state is represented in schema metadata as
required_when.
Runtime Effect Notes
- Route matching and auth-mode enforcement occur at the gateway before forwarding.
- For protected routes, failed JWT/session checks reject the request before service execution.
- For internal calls, services verify short-lived S2S assertions to enforce caller/target trust constraints.
Edge Proxy Reference
Custom Edge Implementations
You can provide your own Edge Proxy implementation, as long as it satisfies the platform contract expected by deployment and runtime flows.
Deployment Assumptions
A compatible Edge Proxy should:
- Sit in front of all public API ingress.
- Require an API key source and forward the resolved key to the Gateway.
- Forward requests to the Z-Arch Gateway endpoint.
- Preserve request method, path, and required auth headers.
- Support forwarding cookies for authenticated browser flows.
- Prevent direct bypass patterns where Gateway endpoints are exposed outside intended routing.
Default Edge Implementation
The default Z-Arch Edge Proxy is a Cloudflare Worker that provides the controlled public entrypoint for live application traffic. It resolves the request API key, removes the edge-only client path segment when applicable, selects a configured Z-Arch Gateway, attaches Google OIDC for Cloud Run, and forwards the request.
The Worker is stateless except for the Cloudflare Cache API entry used to reuse short-lived Google OIDC tokens. Cloudflare supplies network-level DDoS protection. Application-specific traffic controls and advanced balancing policies can be implemented directly in the project’s open-source Worker.
Request Flow
- Parse the incoming
/api/*path. - Resolve the API key from a configured client segment or the incoming
x-api-keyheader. - Strip the client segment only when the Worker injects that client’s configured API key.
- Select a gateway using
DIRECT,RANDOM, orREGION. - Mint or reuse a cached Google OIDC token for the selected gateway audience.
- Forward the request while preserving its method, body, headers, and query string.
zarch.yaml Configuration
edge.load_balancer.method is required when the load_balancer block is present.
edge:
load_balancer:
method: DIRECT # DIRECT | RANDOM | REGION
rate_limiter: nullREGION additionally requires a non-empty map from Cloudflare colo or country codes to deployed GCP regions:
edge:
load_balancer:
method: REGION
region_map:
IAD: us-east1
US: us-east1
DE: europe-west1
DEFAULT: us-east1
rate_limiter: nullThe CLI resolves GCP region names to deployed Gateway URLs before configuring the Worker. The rate_limiter field is reserved and must be null when present.
Worker Runtime Variables
The CLI manages these values as encrypted Worker secrets. They are runtime details, not zarch.yaml field names.
| Variable | Purpose |
|---|---|
GW_URL | Primary Gateway URL used by DIRECT and as the final fallback for other modes |
LB_METHOD | Gateway selection mode: DIRECT, RANDOM, or REGION |
LB_TARGETS | Comma-separated deployed Gateway URLs used by multi-region selection |
LB_REGION_MAP | JSON map of Cloudflare colo/country codes to Gateway URLs for REGION |
API_KEY_<CLIENT> | API key made available when that client has add_to_edge: true |
SA_KEY_JSON | Google service-account key used to mint OIDC tokens for Cloud Run |
Gateway Selection
| Mode | Behavior |
|---|---|
DIRECT | Selects GW_URL. |
RANDOM | Selects a random entry from LB_TARGETS, falling back to GW_URL when only one Gateway is configured. |
REGION | Checks Cloudflare colo, then cf-ipcountry, then DEFAULT, then the first target, and finally GW_URL. |
All three modes use only static configuration supplied during deployment.
API Keys and Edge Paths
Each client in zarch.yaml has an API key stored in Secret Manager. A key is exposed to the Worker only when that client sets add_to_edge: true.
For an injected client key, the first segment after /api identifies the client and exists only at the edge:
| Request | API-key source | Forwarded path |
|---|---|---|
/api/web/v1/session?next=home | API_KEY_WEB | /v1/session?next=home |
/api/v1/session?next=home | Incoming x-api-key | /v1/session?next=home |
If the candidate client segment does not match a configured Worker key, the request must provide x-api-key itself and the path is forwarded without removing that segment. The Z-Arch Gateway performs the authoritative API-key validation.
Google OIDC
The Worker uses SA_KEY_JSON to mint a Google OIDC token whose audience is the selected Gateway origin. It places the token in X-Serverless-Authorization before forwarding to Cloud Run.
Tokens are cached by service-account identity and Gateway audience with an expiration safety window. Cache failures do not introduce persistent state: the Worker can mint a fresh token and continue.
Custom Implementations
The Worker source and wrangler.jsonc belong to the application project and may be extended with custom routing, balancing, observability, or traffic-control behavior. Z-Arch maintains the required deployment fields and runtime secrets while preserving additional JSONC configuration.
Runtime Library
Z-Arch Runtime Library (zarch) provides the authentication primitives (ZArchAuth) used by services running within the Z-Arch architecture. It exposes a stable Python API for minting and handling encrypted session cookies, service-to-service trust, and the ZArchExtension interface with lifecycle hooks used during bootstrap and deployment workflows with the Z-Arch CLI.
Quick Start: ZArchAuth
Create a session service with ZArchAuth:
from zarch import ZArchAuth
auth = ZArchAuth()
# Session service entrypoint.
# In a standard Z-Arch deployment, the gateway validates JWT + session cookie
# before protected traffic reaches your business services.
app = auth.session.start()ZArchAuth.session.start()starts an asynchronous Quart server- runs the session service endpoints (
/session,/session/login,/session/logout,/session/verify) - adds optional session hooks for revocation, backend session control, additional validation requirements, and
/sessionresponse enrichment
Common deployment pattern:
- Keep the session service separate from business services.
- Let Z-Arch Gateway enforce end-user auth; app services focus on business logic.
- Use
await ZArchAuth.s2s.sign(...)andZArchAuth.s2s.verify(...)for internal service trust.
ZArchAuth API Reference
ZArchAuth exposes authentication functionality through two namespaces:
from zarch import ZArchAuth
auth = ZArchAuth()
auth.session
auth.s2sauth.session.start()
ZArchAuth.session.start() starts an asynchronous Quart server and returns its application instance.
app = auth.session.start()The returned application exposes the standard Z-Arch session endpoints:
GET /sessionPOST /session/loginPOST /session/logoutPOST /session/verifyGET /session/health
Returns
| Type | Description |
|---|---|
quart.Quart | Initialized asynchronous session service application. |
Session hook decorators
Register hooks with @auth.session.on_login, @auth.session.on_logout, @auth.session.on_verify, or @auth.session.on_session. Hook functions may be synchronous or asynchronous and only need to declare the available named parameters they use.
@auth.session.on_verify
async def verify(sid, uid):
return await session_store.is_valid(sid, uid)Asynchronous hooks are awaited. Synchronous hooks run in a worker thread to avoid blocking Quart’s event loop. A hook accepting **kwargs receives all parameters available for that event.
| Decorator | Available parameters | Return value |
|---|---|---|
auth.session.on_login | sid, uid, tenant, iat, exp | Ignored |
auth.session.on_logout | sid, uid, tenant | Ignored |
auth.session.on_verify | sid, uid, tenant, iat, exp | True when valid; False denies |
auth.session.on_session | uid, email, tenant, sid, iat, exp, claims | dict to merge into the response, or None |
When on_verify is registered, it must return exactly True or False; any other value raises TypeError. When it is not registered, verification uses the stateless fallback and allows a cryptographically valid session.
Parameter meanings:
| Parameter | Description |
|---|---|
sid | Session identifier. For on_session, this is None when no valid session cookie is available. |
uid | Authenticated user identifier. |
tenant | Tenant identifier, or None. |
iat | Session issued-at Unix timestamp. For on_session, this may be None. |
exp | Session expiration Unix timestamp. For on_session, this may be None. |
email | User email address, or None; available to on_session only. |
claims | Decoded user-info claims; available to on_session only. |
auth.session.register_hook(hook_name, callback)
Programmatically register a synchronous or asynchronous hook. This API remains available as an alternative to decorators and applies the same named-parameter filtering.
| Parameter | Type | Description |
|---|---|---|
hook_name | str | One of on_login, on_logout, on_verify, or on_session. |
callback | callable | Hook function; it may declare any subset of the selected hook’s available named parameters. |
auth.session.verify(cookie_val, expect_uid=None, expect_tenant=None, stateful_session=None, session_service_url=None)
Asynchronously decrypt and validate a Z-Arch encrypted session cookie. This method must be awaited.
In standard Z-Arch deployments, cryptographic session cookie verification is performed at the gateway. When stateful verification is enabled, the session service is also queried to determine whether the session remains valid via the on_verify hook. Only use ZArchAuth.session.verify(...) directly in application code for non-standard topologies. Normal use cases should never need to verify session cookies manually.
claims = await auth.session.verify(
cookie_val,
expect_uid="123",
)Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
cookie_val | str | Yes | Raw encrypted Z-Arch session cookie value. |
expect_uid | str | None | No | Expected user ID. Verification fails if provided and it does not match the cookie. |
expect_tenant | str | None | No | Expected tenant ID. Verification fails if provided and it does not match the cookie. |
stateful_session | str | None | No | Set to "TRUE" to perform authoritative state verification through the session service. |
session_service_url | str | None | Conditional | Base URL of the session service. Required when stateful_session="TRUE". |
Returns
A dict containing the verified session payload, including:
{
"sid": ...,
"uid": ...,
"tenant": ...,
"iat": ...,
"exp": ...,
"v": ...
}Raises
ValueError if the cookie is missing, invalid, tampered with, expired, does not satisfy expected identity constraints, or fails stateful verification.
auth.s2s.sign(req, target, url="")
Asynchronously authorize an outbound request to another Z-Arch service. This method must be awaited.
The method adds a short-lived Z-Arch S2S JWT to the request as:
x-zarch-s2s-token: <token>On GCP, when url is supplied, it also adds a Google OIDC identity token using:
X-Serverless-Authorization: Bearer <token>await auth.s2s.sign(
req,
target="orders",
url="https://orders-abc-uc.a.run.app",
)Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
req | object | Yes | Outbound HTTP request object. It must expose a mutable headers attribute supporting header assignment with req.headers[name] = value. |
target | str | Yes | Z-Arch service ID of the intended downstream service. The target must be permitted by this service’s build-derived S2S policy. |
url | str | No | Base URL of the upstream service. On GCP, this is used as the audience when minting the Google OIDC identity token. |
The req parameter is intentionally framework and HTTP-client agnostic. Any request-like object satisfying the documented headers interface may be used.
Raises
PermissionError if the requested target is not authorized by the service’s S2S_ALLOWED_TARGETS policy.
auth.s2s.verify(req)
Verify the Z-Arch S2S credentials on an inbound request.
The method reads the x-zarch-s2s-token request header and validates the token’s signature, issuer, audience, expiration, freshness, and deployment-derived trust policy.
claims = auth.s2s.verify(request)Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
req | object | Yes | Inbound HTTP request object. It must expose a headers attribute supporting req.headers.get(name, default). |
The request object does not need to be a Quart request. Any framework or request implementation exposing the required headers interface is compatible.
Examples include request objects from Quart, FastAPI/Starlette, and other libraries with equivalent header access semantics.
Returns
A dict containing the verified S2S JWT claims:
{
"iss": ...,
"aud": ...,
"iat": ...,
"exp": ...,
"typ": "zarch/s2s",
"v": 1,
}Raises
PermissionError when the S2S token is missing, malformed, expired, signed by an unknown or invalid issuer, addressed to the wrong service, or outside the allowed freshness bounds.
Usage & Examples
Session
Session cookies are stateless by default. If no hooks are registered, cryptographic cookie validation is enough and /session/verify defaults to valid after payload checks.
To enable stateful behavior (revocation, server-side deny lists, tenant-specific controls) and enrich the /session response, decorate the hooks needed by your session service. Each hook may declare only the available parameters it uses.
Any dict returned by on_session is merged after the default uid/email/tenant fields, so matching keys override the built-in values.
For on_session, sid, iat, and exp come from the current encrypted session cookie when it is available and valid; otherwise they are None.
Real-world pattern:
- hash
sidbefore storage - persist session records on login
- mark
revoked_aton logout (idempotent) - deny in
on_verifywhen revoked, missing, or expired - decorate
/sessionwith app-specific profile metadata when needed
from zarch import ZArchAuth
from google.cloud import firestore
from datetime import datetime, timezone
import hashlib
import time
auth = ZArchAuth()
db = firestore.AsyncClient()
def _hash_sid(sid: str) -> str:
return hashlib.sha256(sid.encode()).hexdigest()
@auth.session.on_login
async def on_login(sid: str, uid: str, tenant: str | None, iat: int, exp: int) -> None:
await db.collection("zarch_sessions").document(_hash_sid(sid)).set({
"uid": uid,
"tenant": tenant,
"created_at": datetime.fromtimestamp(iat, tz=timezone.utc),
"expires_at": datetime.fromtimestamp(exp, tz=timezone.utc),
"revoked_at": None,
})
@auth.session.on_logout
async def on_logout(sid: str) -> None:
await db.collection("zarch_sessions").document(_hash_sid(sid)).set({
"revoked_at": datetime.now(tz=timezone.utc),
}, merge=True)
@auth.session.on_verify
async def on_verify(sid: str) -> bool:
doc = await db.collection("zarch_sessions").document(_hash_sid(sid)).get()
if not doc.exists:
return False
data = doc.to_dict()
if data.get("revoked_at") is not None:
return False
expires_at = data.get("expires_at")
return bool(expires_at and expires_at.timestamp() >= time.time())
@auth.session.on_session
async def on_session(
uid: str,
email: str | None,
sid: str | None,
claims: dict,
) -> dict | None:
return {
"display_name": claims.get("name") or email or uid,
"has_active_session": bool(sid),
}
app = auth.session.start()S2S
ZArchAuth.s2s exists so internal calls are explicitly authorized at the application-policy layer, not just network-reachable.
When you await auth.s2s.sign(req, target, url=...):
- Z-Arch always adds
x-zarch-s2s-token: <jwt>(short-lived Ed25519 JWT withiss,aud,iat,exp,typ). - On GCP (
ZARCH_PLATFORM=gcp) and whenurlis provided, it also addsX-Serverless-Authorization: Bearer <google-id-token>.
Why both are used in GCP deployments:
- The Google ID token is the stronger platform authentication mechanism (Cloud Run/IAM identity boundary).
- The Z-Arch token is a service-authorization policy mechanism (enforces caller identity, audience, and Z-Arch trust graph rules).
- Using both gives layered control: Google proves caller identity to the platform, Z-Arch enforces project policy at the service layer.
The Z-Arch token can also be used by itself in non-GCP or non-IAM topologies. In that mode, services still get signed caller identity and audience/policy checks without requiring Google-authenticated services.
Request Object Compatibility:
ZArchAuth.s2s is framework-agnostic and accepts request-like objects through a minimal interface:
await auth.s2s.sign(req, ...)requires an object with a mutableheadersattribute supporting header assignment.auth.s2s.verify(req)requires an object with aheadersattribute supporting.get(name, default).
This allows S2S authentication to work with request objects from Quart, FastAPI/Starlette, httpx, and other libraries that expose compatible headers.
Caller example (adds both headers on GCP):
import httpx
from zarch import ZArchAuth
auth = ZArchAuth()
async def call_orders_service() -> dict:
req = httpx.Request(
"POST",
"https://orders-abc-uc.a.run.app/internal/create",
json={"sku": "A-100", "qty": 1},
)
# Always injects x-zarch-s2s-token.
# On GCP + url provided, also injects X-Serverless-Authorization: Bearer <google-id-token>.
await auth.s2s.sign(req, target="orders", url="https://orders-abc-uc.a.run.app")
async with httpx.AsyncClient(timeout=30) as client:
response = await client.send(req)
response.raise_for_status()
return response.json()Receiver example (verifies Z-Arch policy token):
from quart import Quart, jsonify, request
from zarch import ZArchAuth
app = Quart(__name__)
auth = ZArchAuth()
@app.post("/internal/create")
async def create_order():
# Verifies x-zarch-s2s-token signature, audience, issuer trust, and freshness.
try:
claims = auth.s2s.verify(request)
except PermissionError as e:
return jsonify({{"error": f"S2S token invalid: {{e}}"}}), 401
caller_service = claims["iss"]
return jsonify({"ok": True, "caller": caller_service}), 200S2S verification data is deployment-derived and local at runtime:
SERVICE_ID: current service identity.S2S_PUBLIC_KEYS_JSON: trusted caller public keys by service ID.S2S_ALLOWED_TARGETS: mint-time policy for where this service may call.
Security Model Summary
- Session cookies are encrypted and stateless by default; stateful controls are added explicitly via
on_login,on_logout, andon_verifyhooks, while/sessionresponse enrichment is opt-in viaon_session. - In the intended Z-Arch platform flow, gateway/session components handle end-user auth so application services do not need to implement cookie auth logic directly.
ZArchAuth.s2s.sign(...)andZArchAuth.s2s.verify(...)enforce short-lived signed service-to-service trust with explicit caller/target validation.- Auth helpers fail closed: invalid, expired, tampered, or unauthorized credentials raise errors that should map to
401/403. - Secret material (cookie encryption keys, S2S keys, API credentials) should come from secure secret management and never be hardcoded or logged.
Extension Interface
This document describes the extension-facing interface exposed to extensions through the project_context argument passed into lifecycle hooks. This is the stable API extensions should use. It is intentionally narrow, safe, and versioned by Z-Arch.
If you are authoring an extension, you should only access functionality via project_context (not internal modules).
Quick Start
A minimal extension looks like:
from typing import Any, Dict
from zarch.extensions.base import ZArchExtension
class Extension(ZArchExtension):
async def post_service_deploy(self, project_context, extension_configuration: Dict[str, Any]) -> None:
project_context.log("Hello from my extension!")The project_context object is your primary tool. It provides:
- project metadata (project ID, region, repo path)
- config accessors
- safe prompt helpers
- GCP helpers (secrets, service URLs, env vars, service accounts)
- Cloudflare helpers and a local-only GitHub helper
Lifecycle hooks are async. Await async project_context helpers such as gcloud, Secret Manager, Cloud Run service metadata, connected-repo lookup, and Cloudflare operations. Logging, config accessors, prompts, and github() are synchronous. Remote extensions do not receive the human GitHub token: project_context.github() and GitHub API mutation require local execution or a separately designed machine credential.
Z-Arch discovers extensions through their zarch.extensions Python entry point.
For example, this entry-point declaration makes firestore-sessions the canonical identifier used by extension:
[project.entry-points."zarch.extensions"]
firestore-sessions = "zarch_ext_firestore_sessions:Extension"Lifecycle Hooks (from ZArchExtension)
Extensions can implement any subset of these methods. Hook methods must be defined with async def; Z-Arch rejects synchronous hook implementations. Each hook receives project_context and the extension-specific configuration block.
pre_project_bootstrap(project_context, extension_configuration)
Runs before initial project bootstrap, but after prompting and repo cloning.
Example
async def pre_project_bootstrap(self, project_context, extension_configuration):
project_context.log("Preparing custom bootstrap")post_project_bootstrap(project_context, extension_configuration)
Runs after initial project bootstrap.
Example
async def post_project_bootstrap(self, project_context, extension_configuration):
domain = project_context.config_get("domain")
project_context.log(f"Project domain is {domain}")pre_service_deploy(project_context, extension_configuration)
Runs before a Cloud Run service is deployed.
Example
async def pre_service_deploy(self, project_context, extension_configuration):
project_context.log("Preparing service deployment")post_service_ensureSA(project_context, extension_configuration)
Runs immediately after the service runtime service account has been ensured/created.
Example
async def post_service_ensureSA(self, project_context, extension_configuration):
event = project_context.get_event_data() or {}
sa = ((event.get("payload") or {}).get("service_account") or {}).get("email")
project_context.log(f"Service SA ready: {sa}")post_service_deploy(project_context, extension_configuration)
Runs after a Cloud Run service has been deployed.
Example
async def post_service_deploy(self, project_context, extension_configuration):
project_context.log("Service deployed successfully")pre_gateway_deploy(project_context, extension_configuration)
Runs before the Z-Arch gateway is deployed.
Example
async def pre_gateway_deploy(self, project_context, extension_configuration):
project_context.log("Preparing gateway deployment")post_gateway_ensureSA(project_context, extension_configuration)
Runs immediately after the gateway service account has been ensured/created.
Example
async def post_gateway_ensureSA(self, project_context, extension_configuration):
payload = (project_context.get_event_data() or {}).get("payload") or {}
project_context.log(f"Gateway SA: {payload.get('service_account', {}).get('email')}")post_gateway_deploy(project_context, extension_configuration)
Runs after the Z-Arch gateway has been deployed.
Example
async def post_gateway_deploy(self, project_context, extension_configuration):
project_context.log("Gateway deployed successfully")pre_job_deploy(project_context, extension_configuration)
Runs before a Cloud Run job is deployed.
Example
async def pre_job_deploy(self, project_context, extension_configuration):
project_context.log("Preparing job deployment")post_job_ensureSA(project_context, extension_configuration)
Runs immediately after the job runtime service account has been ensured/created.
Example
async def post_job_ensureSA(self, project_context, extension_configuration):
payload = (project_context.get_event_data() or {}).get("payload") or {}
project_context.log(f"Job SA: {payload.get('service_account', {}).get('email')}")post_job_deploy(project_context, extension_configuration)
Runs after a Cloud Run job has been deployed.
Example
async def post_job_deploy(self, project_context, extension_configuration):
job_id = project_context.config_get("jobs[0].id")
project_context.log(f"Job {job_id} deployed successfully")pre_scheduler_deploy(project_context, extension_configuration)
Runs before a Cloud Scheduler job is deployed.
Example
async def pre_scheduler_deploy(self, project_context, extension_configuration):
project_context.log("Preparing scheduler deployment")post_scheduler_ensureSA(project_context, extension_configuration)
Runs immediately after the scheduler service account has been ensured/created.
Example
async def post_scheduler_ensureSA(self, project_context, extension_configuration):
payload = (project_context.get_event_data() or {}).get("payload") or {}
principal = payload.get("principal", {}).get("id")
project_context.log(f"Scheduler principal with SA ready: {principal}")post_scheduler_deploy(project_context, extension_configuration)
Runs after a Cloud Scheduler job has been deployed.
Example
async def post_scheduler_deploy(self, project_context, extension_configuration):
scheduler_id = project_context.config_get("schedulers[0].id")
project_context.log(f"Scheduler {scheduler_id} deployed successfully")pre_topic_deploy(project_context, extension_configuration)
Runs before a Pub/Sub topic is deployed.
Example
async def pre_topic_deploy(self, project_context, extension_configuration):
project_context.log("Preparing topic deployment")post_topic_deploy(project_context, extension_configuration)
Runs after a Pub/Sub topic has been deployed.
Example
async def post_topic_deploy(self, project_context, extension_configuration):
topic_id = project_context.config_get("topics[0].id")
project_context.log(f"Topic {topic_id} deployed successfully")Hook Payload Matrix
Lifecycle event payloads are additive schema-v1 summaries. Use .get(...) and tolerate unknown keys.
| Hook | Key payload fields (summary) |
|---|---|
pre_project_bootstrap | project_id, module, principal, repo, create_gcp_project, regions, domain, edge_proxy, userbase |
post_project_bootstrap | status, repo, domain, edge_proxy, userbase, clients, control_plane_ready, gateway_deployed |
pre_service_deploy | principal, resource_type, source, endpoint, flags, routes, targets, env, schema, control_plane_args (wrapper) |
post_service_ensureSA | principal, service_account, resource_type, source, endpoint, flags, targets, routes, env, schema |
post_service_deploy | deployment, inbound_callers, outbound_targets, s2s, env, endpoint, targets |
pre_gateway_deploy | principal, rotate_session_key, min_instance, auth_profile, session, trial_mode, control_plane_args (wrapper) |
post_gateway_ensureSA | principal, service_account, rotate_session_key, min_instance, auth_profile, session, trial_mode |
post_gateway_deploy | deployment, gateway, session, s2s, env |
pre_job_deploy | principal, source, flags, targets, env, control_plane_args (wrapper) |
post_job_ensureSA | principal, service_account, source, targets, flags, env |
post_job_deploy | deployment, targets, target_summary, s2s, env |
pre_scheduler_deploy | principal, schedule_mode, schedule, timezone, paused, targets, target_count |
post_scheduler_ensureSA | principal, service_account, schedule_mode, schedule, timezone, paused, targets, target_count |
post_scheduler_deploy | service_account, schedule_mode, schedule, timezone, paused, targets, target_summary, created_scheduler_job_ids |
pre_topic_deploy | principal, subscribers, subscriber_ids, subscriber_count, publisher_candidates, publisher_candidate_count |
post_topic_deploy | principal, subscribers, subscriber_ids, subscriber_count, publishers, publisher_ids, publisher_count |
All payloads avoid secret values (for example: env var values, session keys, gateway URL/suffix secrets).
Manual Hook Triggering
Use the CLI to manually dispatch lifecycle hooks for configured extensions:
zarch ext trigger pre_service_deploy
zarch ext trigger post_service_ensureSA my-extension
zarch ext trigger post_service_deploy my-extension
zarch ext trigger post_gateway_deploy audit cachehook_namemust be one of the lifecycle hooks defined byZArchExtension.extension_blockis an optional argument list. Values must match extension block names underextensions:inzarch.yaml.- Without
extension_blockarguments, all configured extension instances are considered, and each installed extension that implements the selected hook is invoked. - Dispatch follows the normal hook execution policy (
localvsremote) used by live deployments. - Manual dispatches include minimal event metadata where
sourceis"manual"andpayload.extension_nameslists any explicitly selected extension blocks.
Manual triggers may not work for extensions that require live metadata context to function.
project_context Interface
The sections below describe all available attributes and methods exposed to extensions. Use them as the primary API surface.
Core Attributes
These attributes represent the current project state in a safe, read-only form.
project_context.id(str)- The active GCP project ID.
- Example:
"my-gcp-project"
Example
project_id = project_context.id project_context.log(f"Deploying project {project_id}")project_context.region(str)- The active region for this deployment run.
- Example:
"us-east1"
Example
region = project_context.region project_context.log(f"Active region: {region}")project_context.project_root_path(pathlib.Path)- Absolute path to the project root directory.
Example
root = project_context.project_root_path project_context.log(f"Root path: {root}")project_context.non_interactive(bool)- True if Z-Arch is running in non-interactive mode.
Example
if project_context.non_interactive: project_context.log("Running non-interactively")project_context.config(zarch_cli.helpers.config.Config)- The loaded Z-Arch config object.
- Most extensions should use the
config_get,config_set, andconfig_savehelpers instead of accessingconfigdirectly.
Example
cfg = project_context.config project_context.log(f"Config loaded from: {cfg.root}")
Event Metadata
get_event_data() -> dict[str, Any] | None
Read optional metadata for the lifecycle hook currently being dispatched.
- This may be
Nonewhen metadata is unavailable. - Keys are additive and may grow over time; extensions should tolerate unknown keys.
- Known envelope fields include:
schema_version(integer)source("live"or"manual")hook(hook name)timestamp(UTC ISO-8601)resource(e.g. kind/id/region)payload(hook-specific details, may be empty)
Example
event = project_context.get_event_data() or {}
payload = event.get("payload") or {}
principal = payload.get("principal") or {}
service_account = payload.get("service_account") or {}
project_context.log(
f"Hook={event.get('hook')} principal={principal.get('kind')}:{principal.get('id')} "
f"sa={service_account.get('email')}"
)Logging
log(message: str, level: str | None = None) -> None
Write a styled message to the Z-Arch console.
levelis optional and used only to tag the message (e.g. “INFO”, “WARN”).
Example
project_context.log("Preparing extension steps", level="info")Command Execution
async run_command(command_parts: list[str]) -> tuple[str, int]
Run a local shell command. Returns (stdout, exit_code).
Example
out, code = await project_context.run_command(["echo", "hello"])
if code == 0:
project_context.log(out.strip())async gcloud(command_parts: list[str]) -> tuple[str, int]
Run a gcloud command using the embedded or system gcloud binary. Returns (stdout, exit_code).
Example
out, code = await project_context.gcloud(["projects", "list", "--format=value(projectId)"])
if code == 0:
project_context.log("Projects:\n" + out)Config Access
config_get(key: str, default: Any = None) -> Any
Fetch a config value using dotted path notation.
Example
domain = project_context.config_get("domain", "")
project_context.log(f"Domain: {domain}")config_set(key: str, value: Any) -> None
Set a config value in memory (does not write to disk).
Example
project_context.config_set("gateway.session.stateful", False)config_save() -> None
Persist config changes to zarch.yaml.
Example
project_context.config_set("gateway.session.stateful", False)
project_context.config_save()Prompts
These are safe wrappers around Z-Arch’s prompt system.
ask(message: str, default: str | None = None, required: bool = True, validate: Callable | None = None) -> str
Prompt the user for a string value.
Example
name = project_context.ask("What is the service name?", default="session")choice(message: str, choices: list[str], default: str | None = None, sub_prompt: str = "") -> str
Prompt the user to select a single option.
Example
region = project_context.choice("Select region", ["us-east1", "us-west1"], default="us-east1")multichoice(message: str, choices: list[str], default: list[str] | None = None, sub_prompt: str = "(space to toggle, enter to confirm)") -> list[str]
Prompt the user to select multiple options.
Example
features = project_context.multichoice("Enable features", ["cdn", "auth", "logging"])yes_no(message: str, default: bool = True, sub_prompt: str = "") -> bool
Prompt the user for a yes/no response.
Example
confirm = project_context.yes_no("Proceed with cleanup?", default=False)review_and_confirm() -> None
Render the config and ask the user to confirm. Useful before sensitive operations.
Example
project_context.review_and_confirm()GCP Helpers
These helpers wrap common GCP operations and automatically use the project context’s id and region where applicable.
async ensure_service_account(service_account_name: str, display_name: str | None = None) -> str
Ensure a service account exists, creating it if missing.
service_account_namecan be either a full email (name@project.iam.gserviceaccount.com) or just the short name.- If
display_nameis omitted, Z-Arch derives it from the service account name.
Example
sa = await project_context.ensure_service_account("zarch-ext")
project_context.log(f"Service account: {sa}")async secret_exists(secret_name: str) -> bool
Check if a Secret Manager secret exists in the current project.
Example
if not await project_context.secret_exists("my-secret"):
project_context.log("Secret does not exist")async store_secret(secret_name: str, secret_value: str) -> None
Create or update a Secret Manager secret with a new version.
Example
await project_context.store_secret("my-secret", "super-secure-token")async get_secret(secret_name: str) -> str
Fetch the latest version of a Secret Manager secret.
Example
token = await project_context.get_secret("my-secret")async get_service_url(service_name: str) -> str
Fetch the Cloud Run service URL for a named service in the current region.
Example
session_url = await project_context.get_service_url("session")
project_context.log(f"Session URL: {session_url}")async get_env_var(service_name: str, env_var_key: str) -> str
Read a specific environment variable from a deployed service or function.
Example
public_key = await project_context.get_env_var("zarch-gateway", "S2S_PUBLIC_KEY")async set_env_vars(service_name: str, env_vars: dict[str, str]) -> None
Set or update environment variables on a deployed service or function.
Example
await project_context.set_env_vars("session", {"SESSION_TTL": "1209600"})GitHub
github()
Return the locally authenticated GitHub client (PyGitHub-style client used internally by Z-Arch). It is available only for local execution. Remote extension runtimes receive no human GitHub token; use a separately designed machine credential if a remote extension must mutate GitHub.
Example
gh = project_context.github()
user = gh.get_user()
project_context.log(f"GitHub user: {user.login}")async get_connected_repo() -> tuple[str, str]
Return the connected repo fullname and branch as ("owner/repo", "branch").
Example
repo, branch = await project_context.get_connected_repo()
project_context.log(f"Connected repo: {repo} ({branch})")Cloudflare
These helpers manage Cloudflare workers and pages as used by Z-Arch.
async update_edge_proxy(project_name: str | None = None) -> None
Update the edge proxy worker for the project.
- If
project_nameis omitted, it is inferred from the connected repo name.
Example
await project_context.update_edge_proxy()async set_edge_proxy_envs(env_vars: dict[str, str], project_name: str | None = None) -> bool
Set environment variables on the edge proxy worker.
- Returns
Trueon success,Falseon failure.
Example
ok = await project_context.set_edge_proxy_envs({"API_VERSION": "v1"})
if not ok:
project_context.log("Failed to update edge envs", level="warn")async deploy_cf_worker(script_name: str, repo_root_dir: str, repo_full: str | None = None, branch: str | None = None, domain: str | None = None) -> None
Deploy a Cloudflare Worker from the connected repo.
script_name: Worker script identifierrepo_root_dir: Root path in the repo to deployrepo_full: Optionalowner/repooverridebranch: Optional branch overridedomain: Optional custom domain
Example
await project_context.deploy_cf_worker(
script_name="my-worker",
repo_root_dir="services/edge",
branch="main",
)async set_worker_route(script_name: str, domain: str, route: str = "/api/*") -> None
Attach a route to a worker script.
Example
await project_context.set_worker_route("my-worker", "example.com", "/api/*")async deploy_cf_pages(domain: str, project_name: str | None = None, repo_full: str | None = None, branch: str | None = None) -> None
Deploy a Cloudflare Pages project from the connected repo.
Example
await project_context.deploy_cf_pages("example.com")End-to-End Example
A realistic extension that uses multiple helpers:
from typing import Any, Dict
from zarch.extensions.base import ZArchExtension
class Extension(ZArchExtension):
async def post_service_deploy(self, project_context, extension_configuration: Dict[str, Any]) -> None:
project_context.log("Post-deploy hook starting")
# Read config
domain = project_context.config_get("domain", "")
if not domain:
project_context.log("No domain configured", level="warn")
return
# Ensure a secret exists
if not await project_context.secret_exists("edge-api-key"):
await project_context.store_secret("edge-api-key", "replace-me")
# Update edge proxy envs
await project_context.set_edge_proxy_envs({"API_VERSION": "v1"})
# Deploy pages site
await project_context.deploy_cf_pages(domain)
project_context.log("Post-deploy hook complete")zarch.yaml
Every extension block has an instance name (the mapping key) and a required extension identifier field. The identifier must exactly match the extension’s zarch.extensions Python entry-point name.
extensions:
my-extension:
extension: "my-extension"
required_roles: []
config:
example_key: example_valueValues under config are passed to the extension hook at runtime.
Named multiple instances
Use distinct mapping keys to configure the same installed extension more than once. Each block must repeat the same canonical extension identifier and may provide independent configuration:
extensions:
primary-sessions:
extension: "firestore-sessions"
required_roles:
- "roles/datastore.owner"
config:
collection_name: "primary_sessions"
secondary-sessions:
extension: "firestore-sessions"
required_roles:
- "roles/datastore.owner"
config:
collection_name: "secondary_sessions"Both blocks dispatch to the installed firestore-sessions entry point, once per instance, with that instance’s config value. Manual zarch ext trigger selection via extension_block arguments uses the instance name(primary-sessions or secondary-sessions), not the entry-point identifier.
Explicit roles
Include all GCP IAM roles required by the service account that runs an extension in required_roles.
Notes and Best Practices
- Prefer
config_get/config_setover accessingproject_context.configdirectly. - Use
log()for all extension output to stay consistent with Z-Arch UX. - Avoid raw shell calls unless absolutely necessary; use provided helpers first.
- Never log secrets or gateway URL suffixes.
If you need additional helpers, consider filing a request rather than importing internal modules directly.
