Queries throughout the codebase filter on generations.profile_id,
generations.status, generations.created_at, story_items.story_id,
story_items.generation_id, generation_versions.generation_id, and
profile_samples.profile_id with every request. Without indexes SQLite
falls back to a full table scan; as history grows (hundreds or thousands
of generations) these scans become the dominant latency.
Changes:
- Add index=True on the most-queried FK and sort columns in models.py so
new installs get them from Base.metadata.create_all
- Add _migrate_add_indexes() called from run_migrations() so existing
installs get the same indexes on next startup (uses CREATE INDEX IF
NOT EXISTS — idempotent, <10 ms on any realistic dataset)
Switch from the default DELETE/ROLLBACK journal to WAL so concurrent
readers (SSE status polls, history queries) are not blocked while the
generation worker holds a write transaction. Set a 5-second busy
timeout to eliminate "database is locked" errors under brief write
contention.
Both PRAGMAs are applied via a custom creator function so every
connection in the pool gets the settings at open time, not just the
first one.
Python 3.12 deprecates datetime.utcnow() with a DeprecationWarning and
it will be removed in a future release. Replace all call-site usages in
services, routes, and utils with datetime.now(UTC), and replace the
SQLAlchemy ORM column defaults (which used the bare function reference
datetime.utcnow) with lambda: datetime.now(UTC) so that the returned
objects are timezone-aware and consistent with Python best practice.
Affected: database/models.py, services/{stories,history,profiles,
channels,export_import}.py, routes/{profiles,tasks}.py, utils/tasks.py
list_stories() previously executed one COUNT(story_items) query per
story in a Python loop. With N stories that is N+1 round-trips to
SQLite regardless of list length. Replace with a single aggregated
GROUP BY query that fetches all counts at once, then populate each
StoryResponse from a dict lookup.