# Parking Lot Detection — Project Architecture Overview

## Problem and Approach

The goal of this project is to produce a nationwide parking lot segmentation layer from NAIP (National Agriculture Imagery Program) aerial imagery. NAIP tiles are captured at 0.6 meters per pixel in four bands — red, green, blue, and near-infrared — making them well-suited for impervious surface classification. The system takes those raw GeoTIFF tiles, runs a fine-tuned instance of Meta's SAM2 model over them, and outputs georeferenced GeoJSON polygons and MBTiles vector tilesets suitable for map rendering.

The model of choice is SAM2 (Segment Anything Model 2), specifically the `sam2.1_hiera_tiny` architecture for its favorable memory profile relative to accuracy. Rather than treating this as a traditional semantic segmentation problem with a purpose-built encoder, the project exploits SAM2's prompt-conditioned mask decoder and its hierarchical Hiera vision transformer backbone, adapting them to the aerial remote sensing domain through a structured fine-tuning progression.

---

## Data Pipeline

Ground truth labels are derived from OpenStreetMap parking lot polygons, fetched via the OSM map API using a lightweight bounding-box downloader (`download_osm.py`). The OSM geometries are rasterized against the NAIP tiles to produce binary parking masks at full 1024×1024 resolution. A companion road mask layer is also derived from OSM, used during training as a soft spatial prior to suppress false positives on road pixels.

The data is organized in S3 under a tile-grouped structure:

```
chips/{split}/{tile_name}/images/tile_XXXXXX.tif    ← 4-band NAIP
chips/{split}/{tile_name}/labels/tile_XXXXXX.tif    ← binary parking mask
chips/{split}/{tile_name}/roads/tile_XXXXXX.tif     ← binary road mask
chips/{split}/{tile_name}/manifest.json             ← {filename: {has_parking, has_road}}
```

The per-tile `manifest.json` is a critical optimization. Rather than downloading every label TIFF to determine class balance at dataset construction time, the dataset class first paginates the S3 prefix for image keys, groups them by tile, downloads only the tiny manifest for each tile, and uses those flags to make inclusion decisions. This avoids pulling gigabytes of label data just to balance the dataset.

Empty chips (no parking pixels) are capped at a configurable fraction of the positive set — currently 35% — using a ratio formula: `max_empty = n_positive × frac / (1 − frac)`. This prevents the model from collapsing into a trivial all-background predictor while keeping enough negative examples for specificity.

All images are read with `rasterio` directly into float32 NumPy arrays in CHW order, divided by 255, and normalized using ImageNet-standard RGB statistics. For the 4-channel path, the NIR band uses an independently estimated mean of 0.45 and standard deviation of 0.22. The `DataLoader` uses two persistent worker processes with pinned memory for throughput, and a custom `collate_fn` stacks the uniform-size tensors before returning — image `(B, C, 1024, 1024)`, mask `(B, 1, 256, 256)`, road mask `(B, 1, 256, 256)`, and prompt box `(B, 4)`.

---

## Training Progression

The training has iterated through four distinct scripts, each representing a deliberate shift in strategy. Reading them in sequence is instructive.

**`sam_training.py` (s1)** established the baseline. It uses the `base_plus` model variant, freezes the image encoder entirely, and trains only the mask decoder and prompt encoder using SAM2's standard `set_image_batch` API. The prompt strategy mixed point prompts and bounding box prompts 50/50 per sample, and images were passed as raw uint8 RGB numpy arrays and let SAM2's internal transform pipeline handle normalization. This script proved out the data pipeline and the basic fine-tuning loop but was limited to 3-channel input, discarding NIR entirely.

**`sam_training_s2.py` (s2)** made the first major architectural move: all parameters were unfrozen and the model was trained end-to-end on all four NAIP bands. To accommodate NIR, the `patch_embed.proj` convolution was surgically replaced in-place — a `Conv2d(3, C, ...)` swapped for a `Conv2d(4, C, ...)` with the existing RGB weights preserved and the NIR channel initialized by copying the red channel weights. This is a principled choice: NIR and red reflectance are highly correlated over impervious surfaces, so the red channel provides a far better starting point than zero initialization or random noise. At the same time, the script moved away from `set_image_batch` (which is decorated with `@no_grad` and only accepts numpy input) and began calling `predictor.model.forward_image()` and `_prepare_backbone_features()` directly, allowing gradient flow through the encoder. A differential learning rate scheme was introduced: the encoder trains at `DECODER_LR × 0.1` while the decoder and prompt encoder train at full `DECODER_LR`. A road penalty term was also added to the loss, softly penalizing predicted parking probability on pixels where the road mask is active and the label is negative. The s2 script also introduced encoder gradient norm monitoring as a separate wandb metric to watch for encoder instability during early training.

**`sam_training_s3.py` (s3)** pivoted back to a frozen-encoder refinement strategy, loading the best s2 checkpoint and fine-tuning only the decoder and prompt encoder. This addressed a practical concern: having trained the encoder jointly from a pretrained starting point carried risk of disrupting learned low-level visual features. S3 also introduced a more robust checkpoint loading function that introspects the `patch_embed.proj.weight` shape from the state dict to detect whether the checkpoint is already 4-channel, handling the load-then-expand versus expand-then-load ordering correctly. The prompt strategy was simplified further: 80% of training samples use the full-image box `[0, 0, 1023, 1023]` and 20% use a tight bounding box derived from the mask's actual extent, training the model to handle both coarse "where should I look" prompts and tight "here is the object" prompts. This matches the inference-time prompt generation strategy directly.

**`sam_training_s4.py` (s4, current)** consolidates lessons from the entire progression into a single script with a clean two-phase toggle. The key insight driving s4 is that training the encoder jointly from the start, before the decoder has developed any domain-specific representations, is risky and wasteful. The correct sequence is to first bring the decoder into the domain while the encoder provides stable pretrained features, then optionally fine-tune the encoder in a second pass. A single boolean flag, `TRAIN_ENCODER`, controls the entire behavior of the script. When `False`, the image encoder is frozen using `requires_grad = False` on all encoder parameters and `.eval()` mode is set explicitly each epoch to keep BatchNorm running statistics and dropout stable. The optimizer only receives decoder parameter groups. The dataset loads 3-channel RGB and normalizes with ImageNet stats. When `True`, the patch embed is expanded to 4 channels with NIR copied from red, all parameters become trainable, and the optimizer gets four parameter groups — encoder decay, encoder no-decay, decoder decay, decoder no-decay — with independent learning rates. The warmup schedule was also reworked: rather than a hardcoded step count, `WARMUP_FRAC = 0.05` expresses warmup as a percentage of total training steps, which is resolved at runtime once the DataLoader length is known. This makes the warmup invariant to batch size changes.

---

## The Forward Pass

Because the project bypasses SAM2's high-level `set_image_batch` API entirely, it is worth describing the internal forward pass explicitly.

The image batch `(B, C, 1024, 1024)` enters `predictor.model.forward_image()`, which runs the Hiera hierarchical vision transformer. The backbone produces features at multiple spatial scales — for the tiny model, three levels are produced corresponding to roughly 64×64, 32×32, and 16×16 spatial grids. The raw backbone output is then passed to `_prepare_backbone_features()`, which returns these as a list of `vision_feats` tensors in sequence format `(H×W, B, C)` — the standard ViT token sequence layout. These are then reshaped back to spatial maps `(B, C, H, W)` by permuting and viewing against the known feature sizes from `predictor._bb_feat_sizes`. The highest-level feature (coarsest spatial, richest semantic) becomes `image_embed`. The remaining levels become `high_res_feats`.

The prompt encoder takes the box coordinates and produces `sparse_embeddings` and `dense_embeddings`. For the full-image box prompt, the same box `[0, 0, 1023, 1023]` is broadcast across the entire batch, allowing the prompt encoder and mask decoder to run in a single batched call — a meaningful throughput improvement over the per-sample loop used in earlier scripts. The mask decoder receives the image embedding, its positional encoding via `get_dense_pe()`, both sparse and dense prompt embeddings, and the high-resolution feature list, producing logits at 256×256. `multimask_output=False` suppresses the multi-hypothesis output path, returning a single mask per image. The IoU prediction head produces a scalar confidence score alongside each mask, which is trained with an L1 loss against the actual binary IoU between the predicted mask and ground truth.

---

## Loss Function

The combined loss has four components. The sigmoid focal loss (`α=0.60, γ=2.0`) downweights easy negatives — most of the 1024×1024 image is background — and focuses gradient signal on the ambiguous boundary regions. The Dice loss directly optimizes the overlap metric, complementing focal loss which is pixel-wise. Together they address the severe foreground/background imbalance inherent in parking lot segmentation. The IoU loss trains the decoder's confidence head to be a calibrated signal at inference time rather than a free-floating output. Finally, the road penalty term computes `sigmoid(logits) × road_mask × (1 − label)` and adds it to the loss with a weight of 0.25, creating a soft spatial prior that discourages parking predictions on unlabeled road pixels. The loss is computed at 256×256 resolution, matching the decoder's native output scale.

The learning rate schedule is a linear warmup over the first 5% of training steps followed by cosine annealing to `MIN_LR = 1e-7`. The gradient is clipped by L2 norm at 0.5 before the optimizer step, and the entire forward-backward pass runs under `torch.autocast` in `bfloat16`.

---

## Inference Architecture

The inference system in `inference_tile_s3.py` is a two-pass pipeline designed for production use on large GeoTIFF tiles.

In the first pass, the tile is sliced into 1024×1024 chips with a 20% overlap stride. Chips are batched in groups of 30 through a single backbone call — the expensive Hiera encoder is only run once per chip regardless of how many prompts are generated from it. The full-image box `[0, 0, 1023, 1023]` is used as a uniform prompt for all chips, and the decoder produces logits which are upsampled to 1024×1024 with bilinear interpolation and sigmoidized to probabilities. Overlapping chip predictions are accumulated into a full-tile probability map via an accumulator and count array, then averaged. After thresholding at 0.44, small blobs below 500 pixels are discarded and nearby blobs within 7.5 feet (approximately 12 pixels at 0.6 m/px) are morphologically merged.

The second pass uses the pass-1 blob detections as spatial priors. For each chip that intersects at least one detected blob, the blob bounding boxes clipped to chip coordinates become individual prompts. The backbone is run once per chip batch, but the decoder is run once per prompt within each chip with `repeat_image=True`, reusing the single embedding across all prompts for that chip. This is a deliberate architectural choice: the first pass establishes coarse presence, and the second pass refines the mask geometry using tight semantic prompts. For chips with multiple adjacent blobs, this effectively runs multi-instance segmentation using the decoder's prompt conditioning mechanism. Per-chip probability maps from all prompts are combined with a max operation rather than averaging, since each prompt targets a distinct spatial region.

The final output goes through a vectorization step using `rasterio.features.shapes`, followed by a geometry buffer-dissolve in projected coordinates (EPSG:5070, NAD83/Conus Albers for accurate metre-scale buffering), and is then reprojected to WGS84. For batch processing across many tiles, a chunked spatial union strategy divides the bounding box into a 20×20 grid and performs per-cell union operations before a final global union, keeping peak memory proportional to a single grid cell rather than the full dataset. The final output is serialized to GeoJSON, uploaded to S3, and optionally converted to MBTiles via `tippecanoe` for map rendering delivery.

---

## Engineering Notes

Several engineering choices are worth highlighting specifically. The model loading code in `inference_tile_s3_test.py` introspects the checkpoint's `patch_embed.proj.weight` shape before expanding the architecture, making it robust to both 3-channel and 4-channel checkpoints without requiring any out-of-band metadata. The optimizer's `set_to_none=True` flag on `zero_grad` avoids an unnecessary memset and slightly improves GPU throughput. The `persistent_workers=True` DataLoader setting avoids worker process respawning overhead between epochs, which matters significantly when each worker holds a warm S3 boto3 connection. The `GradScaler` for mixed-precision training is kept in sync with the optimizer step, with explicit `unscale_` called before gradient clipping so that the clip threshold applies to the true gradient scale rather than the scaled one.

The two-phase training design in s4 reflects a principled understanding of transfer learning dynamics. The pretrained SAM2 encoder has seen hundreds of millions of natural images and has learned rich hierarchical representations of edges, textures, and structures. Corrupting those representations early by training on the relatively small NAIP dataset risks both instability and catastrophic forgetting. By first letting the decoder adapt its cross-attention and upsampling path to the aerial domain while the encoder provides stable features, the model builds a reliable decoder representation before the encoder is exposed to gradient updates in phase 2.
