### Overview of the artifacts pipeline

The artifacts pipeline is a multi-stage system that takes a natural-language “analysis” request (plus project context) and ends up with a structured analytics artifact that can be rendered as a PowerPoint deck. Conceptually, it:

1. Plans what sections the artifact should contain and which data needs to be fetched (Stage 0.5).
2. Uses that plan and the retrieved data to write rich, text-based analytical sections (Stage 1).
3. Selects and configures visualization components to support those sections, based on a component registry (Stage 2).
4. Persists the resulting artifact and exposes it so it can be rendered as `.pptx` slides.

The core types driving this are defined in `template_artifacts.py`, and the orchestration logic lives in `ArtifactAgent` in `template_route.py`. The final shape is a `SectionsReport` whose `artifact` contains metadata, attributes (scope), and a list of sections; those sections, plus selected visualization components, can then be transformed into slides.

---

### Stage 0.5 – Planning and MCP data retrieval

Stage 0.5 is a dedicated planning stage, not content generation. Its job is to decide:

- Which sections the artifact should contain, and how they should change across turns.
- Which MCP tools need to run, with which argument payloads, so that the data for those sections is available before we try to write them.

This stage is driven by the `HARDCODED_ARTIFACT_STAGE0_5_PROMPT` in `template_route.py` plus the `Stage05Plan` models in `template_artifacts.py`.

#### Inputs

For Stage 0.5, the agent receives:

- The user’s request and conversation context.
- A `TemplateConfig` (from the request) that includes:
  - `Tools`: names of tools intended for this template.
  - `Knowledge`: prompt keys for domain knowledge.
  - `Recommended_Sections`: a template-defined list of sections, each with:
    - Exact `title`.
    - `central_question`.
    - `expected_output`.
    - `component_guidance`.
- A tool catalog built from MCP config (via `build_mcp_tool` and related infrastructure).
- Tenant, project, and MCP context identifiers that must be used in tool arguments.

#### What Stage 0.5 does

The Stage 0.5 system prompt is explicit:

- On the **first turn**, it must:
  - Use **exactly** the template’s `Recommended_Sections` with original names and metadata.
  - Produce a 1:1 mapping between template sections and planned sections.
  - Not invent or merge sections.

- On **follow-up turns**, it must:
  - Mark sections as `ADD`, `MODIFY`, `PRESERVE`, or `DELETE` according to user instructions.
  - Be conservative about `MODIFY` (only when explicitly requested).
  - Only add new sections when the user explicitly asks.

It also has strict **tool planning rules**:

- It returns only MCP tool calls that should actually be executed on the backend.
- Each tool call has:
  - A `tool_name` that must match the catalog.
  - An `arguments_json` string containing a JSON object with correct parameter names.
- It must call every configured tool that is relevant to the template unless a tool is clearly irrelevant.
- It must cover all data needs for all ADD/MODIFY sections.
- It must include tenant- and project-scoped arguments (`tenant_id`, `project_ids`, `context_id`) exactly as provided.

#### Stage 0.5 output: `Stage05Plan`

The model `Stage05Plan`:

- `sections: List[Stage05SectionPlanItem]`
- `tool_calls: List[Stage05ToolCall]`
- `notes: Optional[str]`

`Stage05SectionPlanItem` describes the planned state of each section:

- `title`: exact section name (from template, unless user-added).
- `central_question`: exact from template for template-defined sections.
- `expected_output`: copied from template.
- `component_guidance`: copied from template.
- `change`: `"ADD" | "MODIFY" | "PRESERVE" | "DELETE"`.
- `priority`: optional ordering.
- `required_tools`: which `tool_name`s this section depends on.

`Stage05ToolCall` describes each MCP call to be executed:

- `mcp_server_id` / `server_label`: how to route the call (optional).
- `tool_name`: matches the MCP catalog.
- `arguments_json`: JSON string containing the arguments.
- `target_sections`: titles of sections that need this call’s results.

This `Stage05Plan` is then used by the backend to:

- Execute each `tool_call` in parallel (using `mcp_client_service` and `tool_aggregator_service`).
- Cache the results and associate them with target sections.
- Hand a consistent "data bundle" to the next stage.

At the end of Stage 0.5, you have:

- A normalized, version-aware section plan.
- A concrete list of MCP jobs to run.
- Aggregated data available for Stage 1’s text generation.

---

### Stage 1 – Textual artifact generation (`SectionsReport`)

Stage 1 takes three things:

- The section plan from Stage 0.5.
- The data fetched via MCP tools, organized per section.
- The template’s domain prompts and knowledge.

Its purpose is to **write the analysis**: each section becomes a well-structured text Q&A, and the whole artifact is wrapped in a consistent artifact-level structure.

#### Core models

In `template_artifacts.py`:

- `SectionData`:
  - `central_question`: the core analytic question for this section.
  - `response`: a detailed, multi-paragraph answer:
    - 5–15 sentences in 2–4 paragraphs.
    - Rich with metrics, dates, comparisons.
  - `sources: List[str]`: citations for data used.

- `Metadata`:
  - `title`, `description`, `source` summary, `period`, `lastUpdated`.
  - `component_guidance` from the template.
  - `section_priority`.
  - `change`: `"ADD" | "MODIFY" | "PRESERVE"` – tracks whether the section is new, updated, or unchanged in this turn.

- `SectionComponent`:
  - `format: "text"` – Stage 1 is fixed to text Q&A.
  - `metadata: Metadata`.
  - `data: List[SectionData]` – always exactly one item, by contract.

A Stage 1 section is thus: “one Q&A-style, text-heavy block, plus metadata and citations.”

At the artifact level:

- `ArtifactAttributes`:
  - `primary_company`
  - `industry`
  - `peer_count`
  - `time_period`  
  These four fields define the **scope** of the analysis. They are intentionally constrained, so you can treat them as slide-level context and deck naming.

- `Artifact`:
  - `version: int` – increments with each new query (see `_get_next_version_number` in `template_route.py`).
  - `section_count: int`.
  - `name: str` – artifact title.
  - `description: str` – 1–2 sentence summary.
  - `attributes: ArtifactAttributes`.
  - `sections: List[Component]`, and `Component` is aliased to `SectionComponent` for Stage 1.

- `SectionsReport`:
  - `artifact: Artifact`.

Stage 1 can be executed **in parallel per section** using `SingleSectionOutput`:

- `SingleSectionOutput.section: SectionComponent`.

This allows the pipeline to fan out generation for each section across multiple LLM calls, then reassemble into one `SectionsReport`.

#### Behavior and outputs

Stage 1 merges:

- The section plan from Stage 0.5 (what to ADD/MODIFY/PRESERVE).
- The actual textual analysis per section (Q&A).
- A unified artifact wrapper (version, attributes, name, description).

The output is:

- A `SectionsReport` object (or JSON equivalent) that can be persisted and also used directly for rendering slides.
- Each section has:
  - A clear title and metadata.
  - A dense, executive-level narrative with citations.
- The artifact-level metadata defines how the deck should be labeled and scoped.

From a slides perspective, Stage 1 produces the **narrative content** of the deck: slide titles, bullet-text or paragraph-level content, and an overall summary.

---

### Stage 2 – Visualization & component selection

Stage 2 takes the text artifact produced in Stage 1 plus data (often the same MCP results) and decides which **visual components** to attach to which sections. These are the building blocks that will become charts, tables, and metric cards in the PowerPoint deck.

The exact visualization models are in `template_artifacts.py`:

#### Visualization component models

Base structure:

- `ComponentBase`:
  - `format`: one of `"table"`, `"key_value"`, `"time_series"`, `"list"`, `"hierarchical"`.
  - `metadata: Metadata` – same metadata type as the text sections, reused so titles/descriptions remain consistent.

Concrete visualization components:

- `TableComponent`:
  - `format: "table"`.
  - `data: TableData`:
    - `headers: List[str]`.
    - `rows: List[List[Cell]]`, where each `Cell` has:
      - `value: Union[str, float, int, None]`.
      - `unit`, `type` (e.g., `"currency"`, `"percentage"`, `"text"`).

- `KeyValueComponent`:
  - `format: "key_value"`.
  - `data: List[KeyValueEntry]`, where each `KeyValueEntry` has:
    - `key` (label).
    - `value: MetricValue` (value + unit + type).
    - `description` (optional textual explanation).

- `TimeSeriesComponent`:
  - `format: "time_series"`.
  - `data: Union[List[TimeSeriesPoint], TimeSeriesMulti]`:
    - `TimeSeriesPoint`: single series-style points (timestamp, value, label).
    - `TimeSeriesMulti`: multiple named series over shared timestamps.

- `ListComponent` and `HierarchicalComponent` exist in the model but are commented out from `VisualizationComponent` union in this snapshot; they’re conceptually there for future or parallel support.

`VisualizationComponent` is defined as:

```python
VisualizationComponent = Union[
    TableComponent,
    KeyValueComponent,
    TimeSeriesComponent,
    # ListComponent,
    # HierarchicalComponent,
]
```

At selection level:

- `ComponentSelection`:
  - `component_id`: the ID from the component registry (e.g., `"data_table"`, `"metric_card"`).
  - `format`: `"table"`, `"key_value"`, `"time_series"`, etc.
  - `data`: the actual data payload for that component (chart data, metrics, etc.).
  - `width`: `"full"` or `"half"` – layout hint for slide rendering.
  - `sources`: citations.
  - `reasoning`: short explanation for why this component was chosen.

- `SectionComponentsSelection`:
  - `components: List[ComponentSelection]` for a single section (usually 1–3 components).

- `ComponentSelectionList`:
  - `components: List[ComponentSelection]` – a more generic collection.

#### How Stage 2 uses these

Stage 2 leverages:

- The **component registry** (`component_registry` and related helpers `get_component_schema`, `get_component_guidance` in `component_registry` and `conversation_tools`).
- The data already fetched via MCP tools in Stage 0.5.
- The section narrative and metadata from Stage 1.

The logic is:

1. For each section that should be visualized:
   - Infer which component types are most appropriate (tables for detailed metrics, key-value for KPIs, time-series for trends, etc.).
   - Use the component registry metadata to understand:
     - What the component expects (schema).
     - In what business/technical cases it’s best used.

2. For each chosen component:
   - Populate `ComponentSelection` with:
     - A `component_id` that matches a registry entry.
     - A `format` aligned with the underlying `VisualizationComponent`.
     - A `data` payload shaped according to the registry’s schema.
     - A layout width hint (`"full"` or `"half"`).
     - `sources` and `reasoning`.

3. Attach each `SectionComponentsSelection` back to the corresponding section (or store them alongside sections). Depending on the code path, Stage 2 can either:
   - Create a separate “visualization plan” that the renderer uses.
   - Or embed references to components within the artifact’s sections or metadata.

From a deck perspective, Stage 2 is where you decide:

- “This section gets a full-width table and a half-width metric card.”
- “This other section gets a time-series chart and a small KPI list.”

The end result is a structure that maps sections to one or more visualization components that can be deterministically converted into slide layouts using a templating engine.

---

### Stage 3 – Persistence and versioning (`ArtifactPersistenceService`)

Once text (Stage 1) and visual selections (Stage 2) are computed, the pipeline persists the artifact via `ArtifactPersistenceService` (`artifact_persistence_service.py`).

#### Inputs

`ArtifactPersistenceService.save_artifact` expects:

- `sections_report: dict` – typically a serialized `SectionsReport` (the artifact-level wrapper with `artifact` containing:
  - `name`, `description`, `attributes`, `sections`).
- `user_id`
- `project_ids`
- `tenant_id`
- `template_id`
- `conversation_id`
- Optional `artifact_id` if continuing an existing artifact.

It reads:

- `artifact_source = sections_report["artifact"]`
- `section_data = artifact_source["sections"]`
- `artifact_name`, `artifact_description`, `artifact_attributes`, `artifact_version`.

#### Behavior

1. **Sanitization**: `_sanitize_null_bytes` walks the data structure and removes `\u0000` from all strings to avoid PostgreSQL errors.

2. **Artifact data assembly**: it builds a dictionary with fields like:
   - `name`
   - `description`
   - `attributes`
   - `is_active`
   - `template_id`
   - `tenant_id`
   - `project_ids`
   - `sections` (the structured sections list)
   - `updated_by`
   - `conversation_id`
   - `version` and `artifact_id`, depending on whether this is a new artifact or a new version.

3. **Versioning**:
   - If `artifact_id` is provided and there is an existing artifact:
     - It increments the version and writes a new row as a new version.
   - Else:
     - It generates a new UUID for `artifact_id`, uses `artifact_version` from the sections report, and creates the first version.

4. **Repository calls**:
   - Uses `ArtifactRepository.create` for new artifacts.
   - Uses `ArtifactRepository.update` for new versions.

The persisted artifact record is now the source of truth for downstream consumers, including anything that renders `.pptx` decks.

---

### Rendering as `.pptx` slides

The code you’ve shown doesn’t include the actual slide renderer, but the data model and pipeline are clearly designed around a PowerPoint-style output:

- **Artifact-level metadata** (`name`, `description`, `attributes`) maps naturally to:
  - Title slide (name + description).
  - A context slide summarizing primary company, industry, peer count, and time period.

- **Sections** (`SectionComponent` with `Metadata` + `SectionData`) map to:
  - One or more slides per section:
    - Slide title from `metadata.title`.
    - Subtitle/description from `metadata.description`.
    - Body text from `SectionData.response`, possibly structured into bullets or paragraphs.
    - Footers or notes built from `sources` and `metadata.source`.

- **Visualization components** (`ComponentSelection` and underlying `VisualizationComponent`) map directly to:
  - Charts, tables, and metric cards on the slides:
    - `TableComponent` → table objects on a slide.
    - `KeyValueComponent` → KPI tiles, scorecards, or numbered lists.
    - `TimeSeriesComponent` → line/bar charts with timelines.
    - `width` → layout decisions (full-slide vs half-slide; multiple components on one slide).

Because all models are explicit about:

- Section priorities (`section_priority`, `priority`).
- Component widths and formats.
- Data schemas and units.

a renderer can deterministically walk the `SectionsReport` plus the associated `ComponentSelection` structures and build a deck:

1. Create deck.
2. Add a title slide from `artifact.name` and `artifact.description`.
3. Add a context slide from `artifact.attributes`.
4. For each section:
   - Add a slide or slide group:
     - Title from `section.metadata.title`.
     - Main body from `SectionData.response` summarized or chunked.
     - Sources in the notes or footer.
   - Insert visualization components as charts/tables on those slides, using `ComponentSelection` and `VisualizationComponent` data.

If the user re-runs the pipeline with updated instructions, a new version of the same artifact is persisted (via `artifact_id` + incrementing `version`), and a new deck can be generated with the updated sections and visuals while preserving the version history.

---

### Summary

- **Stage 0.5**: planning + MCP retrieval.
  - Inputs: template config, tools, knowledge, project/tenant context.
  - Output: `Stage05Plan` – a section plan with change tracking plus a concrete MCP tool-call plan.

- **Stage 1**: artifact text generation.
  - Inputs: Stage 0.5 plan, MCP results, prompts/knowledge.
  - Output: `SectionsReport` – an `Artifact` with version, name, description, attributes, and a list of `SectionComponent` text sections.

- **Stage 2**: visualization and component selection.
  - Inputs: Stage 1 artifact, data, component registry.
  - Output: per-section `ComponentSelection` / `SectionComponentsSelection` structures describing which charts/tables/cards to render and how.

- **Stage 3**: persistence and versioning.
  - Inputs: `SectionsReport` (plus context).
  - Output: persisted artifact versions via `ArtifactPersistenceService`, which are then consumed by a pptx renderer.

The final system output, when passed through the rendering layer, is a complete `.pptx` deck: a set of slides combining text analysis and programmatically generated visuals, fully driven by the structured artifact and visualization models produced by this pipeline.