Followup to commit c84f968 (read-boundary dedup) and commit f75d591
(cleanup CLI). The read boundary filters duplicates out of the
in-memory episodeDict and the CLI cleans up historical duplicates
in the DB, but the underlying pathology — duplicate rows being
created in the first place — was still active on every rescan.
Two layered prevention fixes:
1. Schema-level guard: add UNIQUE(series_id, season, episode_number)
to the episodes table. SQLite's CREATE UNIQUE INDEX requires
no existing duplicates, but the cleanup CLI from f75d591 has
already been run (or is a one-shot prerequisite for users on
older DBs). Future duplicate rows are rejected at the DB layer.
2. Write-site guard: SerieScanner.scan_single_series used to
`extend` the in-memory episodeDict on every rescan of a
series already in keyDict — across N rescans, the same missing
list was appended N times, growing the dict with duplicates that
then flowed through _update_series_in_db into the episodes
table. The fix replaces the cache with the latest scan result
instead of extending, and dedupes within a single call as
defense in depth against a buggy upstream loader.
Defensive dedup is layered three deep:
- schema constraint (this commit, primary)
- scan_single_series replace-not-extend (this commit, secondary)
- episodeDict property read-boundary dedup (commit c84f968,
tertiary — covers legacy DBs that predate the constraint)
Tests:
- Updated test_serie_scanner.test_scan_single_series_existing_entry
to assert the new replace-not-merge behavior (the old assertion
encoded the buggy extend behavior).
- New test_serie_scanner_scan_dedup.py covers the regression
directly: two rescans of the same series with the same missing
list must yield a canonical dict, not an accumulated one.
- test_database_models and test_clean_duplicate_episodes_cli now
use a legacy_engine fixture that drops the UNIQUE constraint,
so the duplicate-row scenarios they exercise (the read-boundary
dedup and the cleanup tool, both meant to defend against
pre-migration state) can still be tested under the new schema.
Verified manually: clean_duplicate_episodes --apply on the user's
backup DB still removes all 633 duplicate rows under the new
schema (the CLI doesn't depend on the UNIQUE constraint — it
operates on whatever rows already exist).
The 'episodes added to download queue never shown' symptom has two
layered causes that masked each other:
1. The 'episodes' table has no UNIQUE constraint on
(series_id, season, episode_number), so historical scans can leave
duplicate rows behind. AnimeSeries.episodeDict iterated the
SQLAlchemy 'episodes' relationship without deduping, so the dict
exposed duplicate entries to list_missing() and the queue UI.
The frontend forwarded the duplicated episode list verbatim to
POST /api/queue/add; the backend's pending-episode dedup then
rejected every duplicate as 'already pending' and the user saw
'Skipped 44 duplicate episodes, Added 0'.
2. selection-manager.downloadSelected counted 'episodes.length'
(the input array) instead of data.added_items.length (the
server-confirmed count). Combined with the backend returning
success on an empty add, the user saw a misleading 'Added 44
episode(s)' toast for an empty queue.
Fix: dedupe at the read boundary. The episodeDict property now
filters duplicate (season, episode_number) pairs from both the
DB-loaded relationship and the legacy _episode_dict_cache path.
Existing duplicate rows in the user's DB are inert — the read
filter makes them invisible to the rest of the stack. The
frontend now trusts the server response, logs a console warning
when an input list shrinks to zero added items, and shows an
accurate toast.
Tests:
- TestEpisodeDictDedup class with 4 regression tests covering:
* duplicate relationship rows deduped
* is_downloaded rows still filtered out
* _episode_dict_cache path also deduped (set by scanners/loaders
that may store duplicates)
* dedup is per-(season, ep_num), preserving legitimate
same-ep-num-across-different-seasons entries
Verified manually against the user's backup DB: 'erased' has 44
duplicate rows in the episodes table; episodeDict now returns
{1: [1..12]} instead of {1: [1,1,2,2,3,3,3,3,...]}, matching the
12-episode canonical list.