# Personal Website — rossklein.com

## Overview

Ross built and maintains his own personal portfolio website at **rossklein.com**. The site is fully self-hosted and serves as both a living resume and a demonstration of his engineering capabilities. The most distinctive feature is an **AI chat assistant** embedded directly on the landing page that answers questions about Ross's background, skills, and experience — essentially an interactive, conversational resume powered by a real production backend he wrote himself.

---

## Architecture

The project is split into two main pieces:

### Frontend — `rossklein-vite`
A **Vite + React 18** single-page application (SPA).

- **React Router v6** for client-side routing (`/`, `/Home`, `/Projects`)
- **Framer Motion** for scroll-driven animations and layout transitions on the landing page
- **p5.js** for a custom interactive physics sketch (bouncing boxes) in the hero section — written as a React component using imperative `useEffect` lifecycle management
- **fun note** - read the childhood_and_personal_life.md note to read a story about why there is a p5 sketch in the background.
- **Custom `useChat` hook** — manages WebSocket connection lifecycle, reconnection logic, session token persistence via `localStorage`, and REST calls for history. No chat library — written from scratch.
- **Marked + DOMPurify** for rendering markdown in chat messages safely
- Responsive layout: chat appears as a sidebar on desktop and collapses to a bottom section on mobile (breakpoint at 900px)
- Built with **Vite**, deployed to production at `rossklein.com`

### Backend — `parking-api` (FastAPI)
A **Python FastAPI** application served via **uvicorn**, deployed at `api.rossklein.com`.

- **WebSocket endpoint** (`/chat/ws`) for streaming AI chat
- **REST endpoints** for chat history (`GET /chat/history`), health check (`GET /chat/health`)
- **Google Maps proxy routes** for a separate parking app (geocoding, Places Autocomplete, Place Details) — API key hidden server-side, protected by a header-based API key (`X-API-Key`)
- **CORS configured** for `rossklein.com`, `www.rossklein.com`, `parking.rossklein.com`, and localhost dev ports

---

## AI Chat Feature (Technical Deep Dive)

This is the most technically interesting part of the site. The chat is not a wrapper around a pre-built widget — it's a custom streaming AI assistant built on the **OpenAI Responses API**.

### How it works

1. The user opens the site; the frontend opens a **WebSocket** to `/chat/ws`.
2. The server performs a **session handshake** — it assigns or restores a session token and sends it back as `{"type": "session", "session_token": "..."}`.
3. When the user sends a message, it streams from the OpenAI **Responses API** (`client.responses.stream`) using `gpt-5.4` with `reasoning={"effort": "low", "summary": "auto"}`.
4. The model has access to a **`read_notes` tool** — a function that reads markdown files from a `notes/` directory on disk. This is how the assistant retrieves detailed information about Ross (resume, college stories, personal interests, etc.) without stuffing everything into the context window upfront.
5. Streaming chunks are forwarded to the client in real time via WebSocket messages of type `"chat"`, with `item_id: "complete"` signaling end-of-stream.
6. A **stop mechanism** is supported: the client can send `{"type": "stop"}` mid-generation and the server cancels the in-flight stream.

### Context window management

- A custom `_build_context` function keeps the **last 100 user turns** in the message history sent to the model.
- Tool messages (`function_call` / `function_call_output`) are only retained if fewer than **10 user messages** follow them — avoiding stale tool context polluting long conversations.
- **tiktoken** (`cl100k_base`) is used for approximate token counting, with periodic `usage` WebSocket messages sent to the client every 5 turns.

### Notes system (RAG-lite)

The assistant's knowledge about Ross lives in a directory of markdown files (`notes/`), organized into sections:
- `resume/` — professional background, projects, skills
- `college/` — college experience
- `childhood_and_personal_life/` — personal stories, hobbies, cooking, etc.

At startup, the backend **auto-indexes** all note files and appends the index to the system prompt so the model knows what files exist. The model then calls `read_notes` with specific filenames when it needs information, rather than having everything pre-loaded.

The system prompt instructs the assistant to **lead visitors toward the resume**, handle recruiters naturally, share personal/lighthearted details only when asked, and protect private information gracefully.

### Session handling

- Sessions are **in-memory** (`dict`) with a **1-hour TTL** (`SESSION_TTL_SECONDS = 3600`).
- A `_sweep_sessions` function runs on each new connection to evict expired sessions.
- `localStorage` on the frontend stores the session token so the conversation persists across page refreshes within the TTL window.
- Reconnect logic in `useChat` retries the WebSocket after a 2-second delay on unexpected closure.

---

## Tech Stack Summary

| Layer | Technology |
|-------|-----------|
| Frontend framework | React 18, Vite 4 |
| Routing | React Router v6 |
| Animation | Framer Motion |
| Creative coding | p5.js |
| Chat rendering | Marked, DOMPurify |
| Backend framework | FastAPI, uvicorn |
| AI | OpenAI Responses API (`gpt-5.4`) |
| Token counting | tiktoken |
| HTTP client (backend) | httpx |
| External APIs | Google Maps Geocoding, Google Places v1 |
| Language | Python (backend), JavaScript/JSX (frontend) |

---

## What Makes This Worth Talking About

- **End-to-end ownership** — Ross designed, built, and deployed every layer: frontend SPA, backend API, AI integration, and hosting infrastructure.
- **Custom streaming WebSocket chat** — not a chatbot SDK. Real streaming with stop/cancel support, session management, and reconnection logic all written from scratch.
- **Tool-augmented LLM** — instead of cramming all context into the system prompt, the assistant dynamically reads markdown files on demand. This is a clean demonstration of tool use / RAG-lite architecture.
- **Context window engineering** — explicit logic to manage which messages and tool calls stay in context across long conversations.
- **Production deployment** — the site is live, with separate prod/dev environment configs (`VITE_*` env vars), CORS policies, and API key gating on sensitive routes.
- **Reasoning model integration** — uses the OpenAI Responses API with reasoning summaries, demonstrating familiarity with newer model APIs beyond standard chat completions.
