Here is an AI generated breakdown of the sheets agent from the production codebase:

### High-level role of the grid builder router

The `grid_builder` router is the public HTTP façade for a fairly sophisticated "agentic spreadsheet" system. It exposes a set of FastAPI endpoints that let clients converse with an AI assistant over Server-Sent Events, mutate grid state in real time, inspect and rewind the operation history, manage grid templates, and integrate Anthropic "skills" and MCP tools. Functionally, this module is deliberately thin: it handles authentication, request/response shaping, SSE streaming details, and OpenTelemetry span boundaries, and then delegates all real business logic to `GridBuilderService` and its helper layer under `grid_builder_service_helper`. Architecturally, you can think of it as the boundary where HTTP concerns, tracing, and security are handled, and where a long-lived, Redis-backed grid state machine is orchestrated via a single service abstraction.

### Authentication, tenancy, and service construction

Every non-trivial endpoint in `grid_builder.py` enforces authentication and role-based access control via dependencies. It uses `get_access_token_bearer()` to extract a token payload, then `get_role_checker(SecurityRole.GENERAL, SecurityRole.ADMIN)` to enforce coarse-grained authorization. Instead of threading user and tenant metadata through each handler, the router centralizes that logic in `_get_service(token_details)`. This helper parses `user_id` and `tenant_id` as UUIDs from `token_details["user"]["user_id"]` and `token_details["user"]["tenant_id"]`, and then constructs a `GridBuilderService` instance with those identifiers.

This design is important: the service is multi-tenant and user-scoped. The `GridBuilderService` uses those IDs to scope database reads and writes (`GridTemplateRepository`, `GridRepository`, `GridSheetRepository`), Redis keys (`grid_redis_service`), secret resolution via `secrets_redis_service`, and tenant-specific skills integration via `TenantService`. From the router's perspective, once it has a `GridBuilderService`, all subsequent operations on conversations, grids, templates, and skills inherit correct tenant and user scoping.

### SSE streaming and keepalive strategy

A central concern of the grid builder router is streaming long-running AI interactions over SSE without having intermediaries or browsers close the connection. The `_with_keepalive` helper is the core of this mechanism. It wraps any async generator of SSE chunks and multiplexes it through an internal queue. A background task drains the underlying generator, placing items into the queue; the outer loop does an `asyncio.wait_for(queue.get(), timeout=interval)`. If there's no new data for the configured interval (15 seconds by default), it yields an SSE comment `": keepalive\n\n"`.

This achieves two things: it prevents idle timeouts in proxies and browsers during long LLM turns, and it acts as a "write probe" against the socket. If the client has disappeared, attempting to send the keepalive will fail and cause the underlying server framework to cancel the generator—this is intentionally relied upon by the downstream `GridBuilderService.process_chat` implementation, whose `finally` block needs to run to persist turns and grid state. The router's streaming endpoints (`/converse` and `/regenerate`) always wrap the engine's generator with `_with_keepalive`, ensuring this behavior is consistent regardless of which particular chat flavor is used.

### Tracing integration and span boundaries

The router integrates with OpenTelemetry using the `tracing_config` helpers. Before starting a chat stream, it attempts to restore a **parent trace context** from the conversation's stored metadata. It calls `ConversationSummaryRepository` with an `AsyncSessionLocal`, loads the conversation row for `request.conversation_id`, and looks up `conversation_metadata["trace_context"]`. If present, it uses `extract(stored)` to obtain a parent context.

For each streaming chat endpoint, the router then obtains a tracer via `get_tracer(__name__)` and starts a span around the entire streaming interaction:

```python
with tracer.start_as_current_span("grid.message" or "grid.regenerate", context=parent_context) as span:
    span.set_attribute("conversation_id", str(request.conversation_id))
    ...
```

This effectively treats "a grid chat turn" as a top-level span, with any lower-level spans (LLM calls, DB operations, MCP tool calls) hanging under it. When the stream completes normally, the span status is set to `StatusCode.OK`; on exceptions, it's set to `StatusCode.ERROR` with the error message, and an `ErrorEvent` SSE is emitted downstream.

A notable detail is how client disconnects are handled. When `asyncio.CancelledError` is raised (typically because the HTTP connection was closed), the router explicitly calls `await chat_gen.aclose()` before re-raising. This ensures the generator's `finally` block, which persists state and updates container IDs, executes before the cancellation propagates. The tracing span is also allowed to close cleanly, with proper error status and logs. Overall, the router defines a clear tracing lifecycle: it attempts to restore continuity from prior turns, marks attributes for observability, and ensures closure semantics are robust even under cancellation.

### Chat endpoints: `/converse` and `/regenerate`

The `/converse` endpoint is the primary entry point for interacting with the grid AI. It accepts a `GridChatRequest`, which encapsulates the conversation ID, the user query, and an optional `client_id` used for broadcast filtering. The router constructs `GridBuilderService`, restores trace context as described above, and defines an inner async generator `generate()` that:

1. Starts an OTel span for the turn.
2. Invokes `service.process_chat(request)` to obtain an async generator of SSE strings.
3. Wraps that generator with `_with_keepalive` to insert periodic comments.
4. Before yielding each chunk, checks `await http_request.is_disconnected()` and logs and terminates early if the client has dropped.

Any unhandled exceptions inside the loop set the span to error and result in an SSE-formatted `ErrorEvent` being yielded, so clients receive a structured error event rather than an abrupt socket close.

The `/regenerate` endpoint is architecturally identical but passes `save_history=False` into `process_chat`. This allows the system to regenerate grid content, including emitting and persisting grid operations, without writing a new conversation history turn. This distinction is captured at the service layer, which uses the `save_history` flag to decide whether to call `insert_turn_input` and whether to save the generated turn back into `conversation_summary` history. The endpoint still participates fully in tracing, keepalive injection, and SSE semantics, and its code path mirrors `/converse` almost line-for-line, with a different span name (`"grid.regenerate"`) to distinguish its behavior in telemetry.

### Context resolution and hydration

While the router itself is intentionally minimal, its behavior can only be understood in conjunction with `GridBuilderService`. Every endpoint that operates on a `conversation_id` first calls `service.resolve_context(conversation_id, cache_id=conversation_id)` inside the service. This method is cached for five minutes using `async_ttl_cache`, which is critical for performance because many endpoints (snapshot, ops, undo/redo, events, templates) are often hit repeatedly for the same conversation.

`resolve_context` looks up the `ConversationSummary` from the DB via `ConversationSummaryRepository`, verifies existence and access implicitly through tenant scoping, and extracts the grid ID (`grid_id`) from `conversation_metadata["grid_id"]`. It also extracts the project IDs, project type, existing conversation title, and any file IDs associated with the conversation. All of that is packaged into a `GridContext` Pydantic model, which is then passed downstream to hydration, history, and persistence logic.

Before any operation that needs the in-memory grid state, the service calls `_hydration.ensure_fresh(grid_id_str, context.grid_id, engine)`. This function (in the helper layer) reconciles Redis, database, and in-process engine state to ensure the `GridBuilderEngine` has an up-to-date representation of sheets, operations, and template. Specifically, it compares the last known grid version held in Redis and the DB against the in-memory view, and if it detects a mismatch or sees that the in-memory state has not been populated yet, it reloads sheets, columns, rows, and operation logs into `GridBuilderEngine` and updates `grid_store`. From the router's viewpoint, this ensures that any call to `get_snapshot`, `get_ops`, `apply_op`, or `events_sse` operates against a consistent, hydrated view of the grid.

### Agentic loop: tools, orchestration, and SSE payloads

The `process_chat` method in the service is where the agentic loop is orchestrated. After resolving context and hydrating the grid, it performs all pre-stream DB work in a single `AsyncSessionLocal` block. The following describes each major component of that setup.

**System prompt and configuration prompts**

`_load_system_prompt` resolves the base system prompt in a strict priority order. If the grid's `GridTemplate` specifies a prompt name under `prompts["system_prompt"]`, that logical name is resolved via `get_prompt_content` against the tenant's prompt store. If not, the service falls back to the tenant-level constant `GRID_BUILDER_SYSTEM_PROMPT_KEY`, again resolved through `get_prompt_content`. This means that both per-template and global behavior can be changed at runtime through the config service, without redeploying code. The resolved string is passed into the agent loop as `system_prompt_base` and forms the foundation for all subsequent instructions, including how the agent should use tools, how it should mutate the grid, and how it should format SSE output.

On top of that, the template's `tool_instruction` and `tool_input_parameters_guidance` fields are part of `GridTemplate` and end up inside the engine's in-memory template store via `engine.update_engine_template`. Those strings are used by the engine when constructing final prompts and tool call descriptions, so they operate as configuration-time tool hints that steer how the model describes, selects, and populates tools. Together, these prompt-based capabilities give you a configurable control plane over the agent's behavior and tool usage strategy.

**Knowledge prompts**

`_load_knowledge` reads `template.knowledge`, which is a list of prompt keys, and resolves each one through `get_prompt_content`. For every key that resolves successfully, it produces a `{key, content}` entry in `knowledge_content`. That entire list is passed into `run_agentic_loop` as `knowledge_content`. The engine treats these entries as a structured knowledge tool: each key names a piece of reusable, tenant-specific knowledge, and the content gives the full text. During planning, the agent can refer to knowledge prompts by key and pull in the content, rather than having the backend decide which prompt to inject for each turn. Because this resolution happens via the prompt store and is cached, you can update knowledge in one place and have those changes propagate to all grids and conversations that rely on the same keys.

**MCP server definitions and tool surface**

`load_mcp_server_configs` reads the template's `connectors_and_tools` configuration from the database and produces a list of "MCP tool configs" that encode which MCP servers and tools should be available. `_build_mcp_servers_list` then converts these configs into concrete Anthropic MCP beta server definitions.

For each config, it resolves the `server_url`, optional `server_label` (used as `name`), and any HTTP headers. If an `Authorization: Bearer ...` header is present, it extracts the token and sets `authorization_token` on the server definition. If no bearer token is found but an `X-API-Key` header is present, it rewrites the URL to include `?x-api-key=...` in the query string. If `allowed_tools` are specified, it adds a `tool_configuration` block that marks the server "enabled" and constrains the agent to that explicit whitelist.

The resulting `mcp_servers_list` is passed into `run_agentic_loop` as `mcp_servers`. From the agent's point of view, this is the complete, authoritative MCP tool surface: each entry represents a server, its authentication, and which tools are legal to call. The backend makes no assumptions about individual tool names beyond this; it simply hands the configuration to the Anthropic client and lets the model choose tools within the allowed set. This is how all MCP-driven capabilities—project queries, data fetches, or platform-specific operations—are surfaced to the agent.

**MCP context creation and scoped instructions**

MCP tools often require contextual parameters, such as which tenant or project they should operate on and which "context" object holds preloaded assets. `_create_mcp_context` encapsulates the creation of that context. It builds a metadata dict containing the current tenant ID and, if present, the active project ID from `GridContext`. It then calls `mcp_client_service.create_context` with that metadata, a TTL, and a "session" scope, using a short-lived `AsyncSessionLocal` for any DB backing that call.

If this call succeeds, the returned `context_id` is fed into `_get_mcp_instructions`, which produces a structured blob of text appended to the system prompt. These instructions explicitly state the tenant ID, the list of project IDs, and the `context_id`, and they tell the model that any MCP tool requiring those parameters must be called with these exact values. The resulting string is passed into the agent loop as `mcp_instructions`. This makes MCP context handling a first-class part of the agent's tool environment: the model is guided to treat `tenant_id`, `project_ids`, and `context_id` as immutable inputs for every tool call, which is key for correctness and multi-tenant safety.

**File context ingestion**

Once `GridContext` is resolved, `process_chat` checks `context.file_ids`. If there are any IDs, it calls `get_files_content` and `format_files_context` from `app.utils.files_context` within the same DB session used to load the template. `get_files_content` fetches file metadata and contents for the given IDs scoped to the tenant, and `format_files_context` produces a single formatted string representation. This string is stored in `files_content_str` and then appended to the system prompt base. There is no separate "file tool" invoked during the loop; instead, the contents are presented as part of the initial context. This avoids additional tool latency for every lookup and ensures that file access respects the existing conversation scoping and prompt caching infrastructure.

**Orchestration and SSE payloads**

With this fully built context, `process_chat` optionally records a new history row via `insert_turn_input`, reuses any existing Anthropic skills container from Redis or conversation metadata, and preloads history messages via `GridHistoryService`, appending the new user message. It then calls `engine.run_agentic_loop(...)`, passing grid ID, messages, conversation and client IDs, knowledge content, MCP servers, the base system prompt, a mid-loop Redis flush callback (`on_state_persisted`), title metadata, container ID, and a per-service-instance `claude_source_id` (`self._stream_id`) used to tag LLM-originated ops. `run_agentic_loop` itself is an async generator that emits SSE strings conforming to the shared `conversation_mode` vocabulary: `text_delta`, `grid_state_update`, `grid_sheet_create`, `grid_template_update`, `completed`, and `error` events. The router simply forwards these chunks over the network, wrapped in keepalive logic.

When the generator completes, or when it's cancelled due to a disconnect, the service's `finally` block launches a background task using `asyncio.create_task` and `asyncio.shield` to persist the "turn": it writes updated op logs, sheets, templates, and MCP container IDs via `GridPersistenceService`. The router's responsibility is to ensure this finally block gets a chance to run; hence the explicit `aclose()` and structured error handling in the `/converse` and `/regenerate` endpoints.

### Direct grid operations: snapshot, ops, apply, undo, redo, sheets

Beyond chat, the router exposes lower-level primitives for manipulating and inspecting grid state. Importantly, both user-initiated operations and agent-initiated mutations use the same underlying op schema and `grid_store` machinery; the only difference is provenance: agent ops carry a `source` field populated from `claude_source_id` (`self._stream_id`), while user ops go through `apply_user_op` directly.

`GET /{conversation_id}/snapshot` delegates to `service.get_snapshot(conversation_id)`, which resolves context, hydrates the engine via `ensure_fresh`, and returns `engine.get_grid_snapshot(grid_id_str)`. This snapshot contains the current columns and rows of the active sheet, reflecting both user and AI-applied operations. During long agent turns, snapshots remain consistent because the `on_state_persisted` callback passed into `run_agentic_loop` invokes `flush_ops_mid_loop`, which writes new operations and updated metadata to Redis so that subsequent `ensure_fresh` calls from any process work from a nearly up-to-date base rather than replaying a long backlog of operations.

`GET /{conversation_id}/ops` returns the operation history via `engine.get_grid_history(grid_id_str)`, again after ensure-fresh hydration. This history is the basis for undo/redo semantics and for audit or replay.

`POST /{conversation_id}/op` is the low-level "apply a user-initiated grid mutation" endpoint. The body is a generic `dict[str, Any]` describing an op such as `SET_CELL` or `ADD_ROW`, plus an optional `client_id` query parameter. The service applies this op via `engine.apply_user_op`, and if the result indicates success, it persists the single op immediately by locating it in the `grid_store`'s op log, calling `save_user_op_turn`, and, if column recipes changed, calling `save_template_only`. This immediate persistence is explicitly engineered to avoid version skew: if a user change were left only in in-memory state, the next `ensure_fresh` call would see a mismatch and might cold-load from the DB, silently discarding the user's mutation.

Undo and redo are handled by `POST /{conversation_id}/undo` and `/redo`, which accept an `UndoRequest` describing mode (`single` vs `stream`), target sequence, stream ID, strategy, and expected sequence. The service resolves context, hydrates, and then calls `engine.undo_ops` or `engine.redo_ops`. When successful, it identifies the relevant op entries in the op log, computes mappings of `reverts_seq` to "reverted updates," and pushes those into `GridPersistenceService.save_reversion_ops_turn`. That method writes reversion ops into persistent storage and synchronizes Redis, ensuring that clients reconciling via `ensure_fresh` see the same logical reversion history the user observed.

Saving full sheet state from the client is supported via `POST /{conversation_id}/sheets`. Here the router accepts a strongly typed `SaveSheetsRequest`, hydrates context and engine, then calls `engine.save_all_sheets`. It persists this bulk sheet state via `GridSheetRepository`—deleting existing sheet records for the grid and reinserting each sheet with its serialized columns and rows—treating the client as the authoritative source of truth. This path is used when the frontend wants to push an authoritative view of all sheets, for example after client-side reordering or heavy edits. Once this completes, the next `get_snapshot` and any export via `download_xlsx`—which calls `engine.get_all_sheets_snapshot` and `generate_grid_xlsx`—will reflect exactly what the client sent.

### Real-time broadcasts: `/events` SSE

The `/events` endpoint is a separate SSE channel designed for broadcast of grid mutations triggered by other clients, including other chat streams. It accepts a `client_id` query parameter used to identify the connecting client. The engine's `events_sse` implementation is expected to subscribe that client to a pub/sub or similar mechanism in Redis and to emit `grid_state_update` events that exclude mutations originating from the same client ID, so the frontend doesn't receive duplicates for its own actions.

The router's role here is minimal: authenticate, authorize, build the `GridBuilderService`, resolve context, hydrate the engine, and then return the result of `engine.events_sse(http_request, grid_id_str, client_id)` as a `StreamingResponse`. It still benefits from the same context resolution and hydration pipeline used by all other operations, so the broadcast semantics are always tenant- and grid-scoped.

### Export and Excel generation

The `/download` endpoint exposes export to Excel. Given a `conversation_id`, the service resolves the grid context, hydrates the engine, synchronously extracts `engine.get_all_sheets_snapshot(grid_id_str)`, and passes that into `generate_grid_xlsx`, which returns a `BytesIO` buffer. The router wraps this in a `StreamingResponse` with an appropriate `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet` media type and `Content-Disposition` header naming the file `grid_{conversation_id}.xlsx`. The design choice to stream from a buffer rather than materializing the entire file in memory per request is pragmatic: the buffer is produced by the helper and can be streamed in chunks, and headers are configured to avoid caching sensitive tenant data.

### Skills integration and template-driven skills filtering

At the tail end of the router, there's a set of endpoints focused on Anthropic "skills" integration. `GET /skills` returns all skills available from the tenant's skills API by delegating to `service.list_skills(session)`. Inside that method, the service obtains tenant-specific credentials via `TenantService.get_ragnarok_config`, constructs a request to the external Ragnarok API (`RAGNAROK_API_SERVER`), and normalizes each returned skill so that `anthropic_skill_id` becomes the canonical `id` used by templates.

`GET /{conversation_id}/skills` narrows this list to those skills enabled for the current grid. `get_grid_skills` resolves `GridContext`, hydrates the engine, fetches the in-memory template via `engine.get_grid_template(grid_id_str)`, and inspects `template.skills`, which is a list of skill IDs. If the template specifies no skills, it returns an empty list immediately. Otherwise, it calls `list_skills` and filters the result down to those entries whose `id` appears in `template.skills`. The returned list is the effective skill set for the grid: the capabilities the agent can legitimately leverage and that the frontend should surface in the UI. This mechanism decouples global skill availability (from the Ragnarok service) from grid-specific skill activation (via templates), and is a good example of cross-cutting concerns—grid templates, external skills APIs, and conversation context—converging in a simple HTTP surface.

### Template lifecycle: CRUD and export

A substantial portion of the router is devoted to grid template administration, scoped by tenant and user but otherwise decoupled from any particular conversation.

`POST /templates` creates a new `GridTemplate` directly from a request body, sets `tenant_id`, `created_by`, and `updated_by` from the service's user and tenant IDs, and persists it with `GridTemplateRepository`. The response returns the Pydantic model's `model_dump()` representation, making it suitable for configuration UIs.

`GET /templates` exposes paginated template listing with filtering options `is_active` and `is_factory_template`, as well as textual search, page, and limit parameters. Behind the scenes, the service simply delegates to `GridTemplateRepository.get_paginated`, passing the tenant ID and filter parameters, then the router maps each item to a dict via `model_dump` before returning the result.

`GET /templates/{template_id}` and `PUT /templates/{template_id}` follow typical REST semantics: fetch and optionally update a specific template by ID, scoping to the tenant. If a requested template does not exist, the router returns a FastAPI `JSONResponse` with a 404 status and a simple `{"detail": "Template not found"}` payload. Delete semantics via `DELETE /templates/{template_id}` perform a hard delete and follow the same 404 behavior when nothing is deleted.

The most interesting template endpoint is `POST /templates/{conversation_id}/export`. It resolves the grid from the conversation context, then delegates to `service.export_grid_template_by_conversation_id`. That method fetches the `Grid` record for the context's `grid_id`, verifies that the grid has a linked `template_id`, and then calls `export_grid_template`, which copies that template and flags the copy as a factory template. The router returns the exported factory template's serialized form or 404 if either the grid or its linked template is missing. This provides a clean promotion path from a working grid configuration, as used by a conversation, into a reusable tenant-level "factory" template.

### In-memory vs persisted template reconciliation

The `PUT /{conversation_id}/template` endpoint drives a more nuanced interaction between in-memory state and persistent templates. When a client sends an `UpdateGridTemplateRequest`, the service resolves context, hydrates the engine, and then—only if `column_recipes` was explicitly present in the request—performs a diff between the existing template loaded from Redis/DB via `_load_grid_template_from_db` and the new recipes.

It resolves an effective column letter for each recipe using `_recipe_letter`. This helper interprets three signals in priority order: a backend-set `column_letter`, a frontend-set `priority` (A/B/C-style values), or a positional index fallback. If the resolved value is numeric, it converts it into a letter via `col_idx_to_letter`. Using that mapping, the service computes two sets: letters that have been removed, and recipes that have been newly added.

For each removed letter, it looks up the current column at that position from `grid_store.get_or_create(grid_id_str).active_sheet.columns`, builds a `DeleteColumnAction` with the column field name, and dispatches it through `grid_store`. For each added recipe, it constructs a corresponding `ColumnDef` and an `AddColumnAction` and dispatches that as well.

After calling `engine.update_grid_template(grid_id_str, update)` to update in-memory state, all these add/remove ops are persisted via `GridPersistenceService.save_reversion_ops_turn` and broadcast to clients via the grid's existing SSE channels. If there are no column changes, the service still synchronizes template and grid state to Redis via `grid_redis_service.set_all`. If there is no DB template yet, it creates a new `GridTemplateDB`, links it to the grid record, and then updates the in-memory Pydantic model with the new ID and tenant ID. The router simply surfaces this behavior as an HTTP `PUT`, but the underlying architecture handles consistent reconciliation among in-memory engine state, Redis caches, and relational DB records.

### History management

History in the grid builder system has two layers: chat history and grid operation history. On the chat side, `GridHistoryService` manages the sequence of messages associated with a conversation. At the beginning of each `process_chat` call, `load_messages` is invoked to retrieve previously stored messages from Redis and/or the database. The service then appends the new user message to that list and uses it as the `messages` argument to `run_agentic_loop`. If `save_history` is true (always the case for `/converse`, never for `/regenerate`), it also inserts a new "input" row in the history table via `insert_turn_input`, capturing the raw user query and obtaining a `history_id`.

At the end of the agentic loop, regardless of whether the stream completed normally, failed, or was cancelled, the service spawns a background task `save_turn` using `asyncio.create_task` and `asyncio.shield`. That task receives the `GridContext`, the `result` dict populated during the loop (containing, among other things, any `container_id` and output metadata), the `history_id`, the engine instance, and the original query. `save_turn` is responsible for writing the assistant's outputs, updated conversation metadata, and any derived state back to the database and Redis. By running this in a shielded `asyncio` task, the service ensures that chat history is durably updated even when the HTTP connection dies mid-stream.

On the grid operation side, `GridPersistenceService` and `grid_store` cooperate to maintain an append-only log of operations. Every mutation—whether user- or agent-initiated—becomes an op record in the in-memory log with a unique sequence number (`seq`) and a `source` field indicating provenance. User ops are flushed immediately via `save_user_op_turn`. Agent ops are flushed intermittently during a turn via `flush_ops_mid_loop` (invoked by the `on_state_persisted` callback) and at the end of the turn via `save_turn`. Undo and redo operations add further "reversion" ops to this log with `reverts_seq` fields marking which original ops they target, and these are persisted via `save_reversion_ops_turn` along with a `reverted_updates` map that lets the persistence layer track which operations are currently active. This unified history model provides both a complete audit trail and the mechanical basis for undo/redo semantics.

### Summary

Taken together, `app/api/routers/grid_builder.py` is a textbook example of a clean HTTP boundary over a complex, distributed grid/LLM system. It centralizes authentication, authorization, tracing, SSE plumbing, and error shaping, and then defers to `GridBuilderService` and its helper modules for all domain-specific behavior: resolving conversations into grid contexts, hydrating state across DB and Redis, orchestrating Anthropic calls and MCP tools, persisting operation histories and templates, and integrating tenant-scoped skills. The router's functions are narrow but carefully engineered: they ensure that long-lived SSE connections are robust against idle timeouts, that tracing spans reflect real user turns, that client disconnects do not corrupt state, and that the multi-tenant boundaries are consistently enforced. As such, it forms a coherent, senior-level design where the HTTP surface area is small and composable, while the grid builder engine beneath it can evolve independently.
