Implement optional single-image intake, Signal ingestion, and multimodal vision analysis

- Add image intake service with format validation (JPEG, PNG, WebP) and EXIF/GPS stripping
- Enforce strict single-image rule across Web and Signal attachment channels
- Implement token-optimized vision downscaling and JPEG compression
- Add IMAGE_CONTEXT pipeline stage with OmniRoute vision routing and resilient failover
- Seed and manage versioned idea-image-interpreter prompt in catalog
- Update Web UI with responsive image picker, preview chip, and Visual Context tab
- Add comprehensive automated test suite in test_image_intake.py
- Update README and Labyricorn devlog
This commit is contained in:
2026-08-23 01:40:22 -07:00
parent 94ff1e4408
commit 41e08611c9
30 changed files with 3784 additions and 278 deletions
@@ -0,0 +1,43 @@
_model: devlog-entry
---
schema_version: 1
---
title: Multimodal Intake: Optional Single-Image Reference Ingestion and Vision Analysis
---
date: 2026-08-23
---
author: Labyricorn
---
summary: Implemented optional single-image reference artifact intake across the public web submission form and inbound Signal gateway, supporting EXIF sanitization, token-optimized vision downscaling, and resilient OmniRoute multimodal interpretation.
---
tags: intake, vision, signal-gateway, omniroute, multimodal, privacy
---
source_commit: e28f31fc2ca1a03cc17b67b61e3ee5ac3f6f2365
---
body:
# Multimodal Intake: Single-Image Reference Ingestion & Vision Context
Ideas often originate as sketches on napkins, whiteboard diagrams, UI wireframes, or architecture blueprints. To support visual thinking without introducing cumbersome gallery management or unbounded token costs, we introduced strict, privacy-first single-image reference intake across ThinkStorm's web and Signal channels.
## Key Architectural Highlights
### 1. Strict Single-Image Rule & Safe Sanitization
- Users can attach at most **one** reference image per idea submission (JPEG, PNG, or WebP).
- Submissions via the web form support seamless multipart upload with image preview and removal controls.
- Signal integration processes inbound attachments with robust filtering: non-images are silently ignored, the first valid image is ingested and sanitized, and any extra attachments are ignored.
- Automatically strips EXIF, GPS coordinates, and camera metadata using Pillow while preserving correct visual orientation via `ImageOps.exif_transpose`.
- Persists canonical reference image files under durable artifact paths (`/root/data/artifacts/{idea_id}/IMG-{idea_id}.{ext}`).
### 2. Token-Optimized Vision Downscaling
- Large high-resolution images can consume prohibitive amounts of vision tokens across model grid tiles.
- The pipeline downscales oversized images to a bounded maximum dimension (1536px) while preserving exact aspect ratios, optimizing payload transmission to OmniRoute.
### 3. Dedicated `IMAGE_CONTEXT` Processing Stage & Resilient Failover
- Introduced a dedicated `IMAGE_CONTEXT` pipeline stage powered by the admin-configurable `idea-image-interpreter` prompt catalog entry.
- Transmits OpenAI-compatible multimodal payloads to OmniRoute's `auto/best-vision` policy.
- Built-in heuristic failover: if upstream vision providers decline the request or experience transient outages, a structured fallback is logged in processor provenance without interrupting the idea's progression through downstream research synthesis.
### 4. Dossier & UI Integration
- Canonical `idea.md` and Gitea repository `README.md` files include dedicated `## Reference Image` and `## Image Context` sections.
- The web UI idea dossier renders the reference image with metadata badges (format, dimensions, file size, SHA-256) and provides a dedicated **Visual & Image Context** report tab.