Followup to commit c84f968 (read-boundary dedup) and commit f75d591
(cleanup CLI). The read boundary filters duplicates out of the
in-memory episodeDict and the CLI cleans up historical duplicates
in the DB, but the underlying pathology — duplicate rows being
created in the first place — was still active on every rescan.
Two layered prevention fixes:
1. Schema-level guard: add UNIQUE(series_id, season, episode_number)
to the episodes table. SQLite's CREATE UNIQUE INDEX requires
no existing duplicates, but the cleanup CLI from f75d591 has
already been run (or is a one-shot prerequisite for users on
older DBs). Future duplicate rows are rejected at the DB layer.
2. Write-site guard: SerieScanner.scan_single_series used to
`extend` the in-memory episodeDict on every rescan of a
series already in keyDict — across N rescans, the same missing
list was appended N times, growing the dict with duplicates that
then flowed through _update_series_in_db into the episodes
table. The fix replaces the cache with the latest scan result
instead of extending, and dedupes within a single call as
defense in depth against a buggy upstream loader.
Defensive dedup is layered three deep:
- schema constraint (this commit, primary)
- scan_single_series replace-not-extend (this commit, secondary)
- episodeDict property read-boundary dedup (commit c84f968,
tertiary — covers legacy DBs that predate the constraint)
Tests:
- Updated test_serie_scanner.test_scan_single_series_existing_entry
to assert the new replace-not-merge behavior (the old assertion
encoded the buggy extend behavior).
- New test_serie_scanner_scan_dedup.py covers the regression
directly: two rescans of the same series with the same missing
list must yield a canonical dict, not an accumulated one.
- test_database_models and test_clean_duplicate_episodes_cli now
use a legacy_engine fixture that drops the UNIQUE constraint,
so the duplicate-row scenarios they exercise (the read-boundary
dedup and the cleanup tool, both meant to defend against
pre-migration state) can still be tested under the new schema.
Verified manually: clean_duplicate_episodes --apply on the user's
backup DB still removes all 633 duplicate rows under the new
schema (the CLI doesn't depend on the UNIQUE constraint — it
operates on whatever rows already exist).
The 'episodes' table has no UNIQUE constraint on
(series_id, season, episode_number), so historical scans can leave
duplicate rows behind. Commit c84f968 added a read-boundary dedup
in AnimeSeries.episodeDict so the rest of the stack never sees
duplicates — but the duplicate rows themselves still bloat the DB
and confuse direct SQL queries.
This commit adds a standalone CLI to find and (with --apply)
delete those duplicate rows. The cleanup keeps the lowest 'id'
per tuple (the oldest insert, which is most likely to have
populated title / file_path fields) and is idempotent.
Usage:
python -m src.cli.clean_duplicate_episodes # dry-run report
python -m src.cli.clean_duplicate_episodes --apply # actually delete
python -m src.cli.clean_duplicate_episodes --max-series 5 # limit report
Override the target DB with DATABASE_URL=sqlite:///path/to.db.
Verified against the user's backup DB: 633 duplicate rows across
231 (series, season, episode) tuples removed cleanly, leaving
1221 unique rows. Re-running reports no duplicates. The cleanup
does not affect the read-boundary dedup — both layers are
defensive in depth.