Skip to content

Fix IMDb ID parsing and remove the discontinued SDBits provider - #12327

Merged
medariox merged 2 commits into
pymedusa:developfrom
XxUnkn0wnxX:fix/imdb-id-parsing
Oct 7, 2026
Merged

medariox merged 2 commits into
pymedusa:developfrom
XxUnkn0wnxX:fix/imdb-id-parsing

Conversation

@XxUnkn0wnxX

@XxUnkn0wnxX XxUnkn0wnxX commented Oct 6, 2026 •

Copy link
Copy Markdown

🎯 Why

This fixes a show-add blocker: malformed optional IMDb metadata can prevent affected shows from being added at all. The failure aborts the add task before the show is saved.

The confirmed example is MAGICAL GIRL RAISING PROJECT restart (TVDB 457780): TVDB returned the series successfully, but parsing its IMDb field raised an uncaught exception. Retrying, including after a restart, left the show absent while Medusa itself remained running.

Optional IMDb metadata should be normalized when usable and safely rejected otherwise, so malformed values do not abort an otherwise valid show addition.

This PR also removes the discontinued SDBits provider following feedback from @duramato.

🛠️ Changes

  • Extract complete IMDb title IDs from URL/path components, ignoring query strings and fragments. Preserve supported numeric IDs, leading zeros, and variable-length IDs; safely reject malformed values.
  • Normalize Series fallbacks and title-search results, preserving external-ID precedence and the existing first-ID policy for comma-separated metadata.
  • Guard the remaining consumers of rejected IDs in Series, EZTV, and Trakt, preventing invalid mappings, crashes, and false matches.
  • Remove the discontinued SDBits provider, its regression cases, theme assets, and fixture entry.
  • Add regression coverage for the exact failing metadata, supported formats, malformed inputs, and completion of the add-queue path.

The IMDb fix is limited to Medusa's parser and its immediate callers; SDBits retirement is a separate cleanup. Vendored imdbpie, dependencies, queue implementation, and homepage behavior are unchanged.

🔗 Related caller changes and why they are needed

The parser returns no identifier for malformed values, so the remaining callers guard that result at existing call sites:

Caller Change and reason
Series metadata fallback and title search Normalize the selected ID before requesting IMDb metadata, so malformed raw values cannot bypass the parser; retain external-ID precedence and first-comma-ID handling
Series external-mapping persistence Skip rejected IMDb mappings rather than inserting a NULL external ID into the database
EZTV Check the normalized value before removing the tt prefix; skip its IMDb-only search when the ID is unusable instead of slicing None
Trakt watchlist matching Require a usable IMDb ID before comparison, preventing unrelated shows from matching through None == None

Valid searches, valid IMDb matches, and primary-indexer matching retain their existing behavior. These remain small guards around the parser's rejection result.

🧹 SDBits removal

@duramato reported that SDBits is permanently closed. I removed its provider module, four dedicated regression cases, fixture entry, and three theme icons. This cleanup does not redesign any other provider.

✅ Validation

I verified the changes through regression tests, the earlier live UI add, isolated runtime checks, frontend checks, and fork CI:

Check Result
Live add of TVDB 457780 Show and 12 episodes persisted on dc17f43; historical proof for the earlier candidate
New IMDb regressions 124 tests remain; included in the full backend suite
Full local backend suite 1,528 passed, 1 expected failure; environment notes below
Frontend ShowResults test and snapshot passed; development theme build and Gulp checks passed
Isolated runtime Accepted legacy SDBits settings, listed 56 providers without SDBits, and returned 404 for /providers/sdbits; live database and config hashes were unchanged
Fork CI 28 checks passed for 6aa24e4; validation draft #17 closed afterward

The PR contains two signed commits based on develop: the original IMDb fix and a separate SDBits removal. The remaining IMDb code is unchanged; the SDBits startup checks used isolated data.

⚠️ Existing API-test annotations

I compared the current fork validation run, original upstream PR run, and unchanged upstream develop run at base 660bc2a. All three passed with 91 passing, 0 failing, 0 errors, 14 skipped, while displaying the same three Show folder does not exist. annotations at medusa/tv/series.py:423.

The Dredd add test queues TVDB 80379 without a show directory. Banner, poster, and fanart lookups each raise a directory exception that Medusa catches and logs as a warning with a traceback. GitHub's Python log matcher turns those tracebacks into error annotations; the API assertions still pass.

These annotations predate this patch. The directory-validation, image-cache, and API-test paths responsible are unchanged; their cleanup remains outside this IMDb fix.

🔎 Evidence

Download the before/after logs and console trace.

🔁 Reproduction and original failure
  1. Open Add Shows, search TVDB for MAGICAL GIRL RAISING PROJECT restart, and select ID 457780.
  2. Submit the add request and wait for the queue to process it.
  3. Before the fix, the task stopped after the successful TVDB response and the show never appeared in the library.

TVDB returned this IMDb field for series 457780:

tt6135388/episodes/?season=2&ref_=ttep

The old parser selected the final slash-separated component, split it on tt, and attempted int('ep'):

ValueError: invalid literal for int() with base 10: 'ep'

Adding the show through TVDB stopped before persistence. Four recorded attempts ended at the same point, including attempts after restarting Medusa. The same parser failure was reproduced from upstream develop at 660bc2a.

The fixed parser extracts tt6135388 / 6135388. The regression retains the exact input independently of future TVDB metadata changes. This validates the ID's format, not whether the external mapping points to the correct title.

📋 Successful live add: console, log, and database evidence

I repeated the UI add on dc17f43, on 2026-10-06, AEDT (UTC+11). This is historical runtime proof from the earlier candidate; the IMDb code is unchanged in 6aa24e4.

Browser console excerpts:

15:34:41.539 XHR POST
http://localhost:8081/api/v2/series
[HTTP/1.1 201 Created 3ms]

15:35:27.720 Adding MAGICAL GIRL RAISING PROJECT restart as it wasn't found in the shows array

Application log excerpts:

2026-10-06 15:34:43 INFO     SHOWQUEUE-ADD :: [dc17f43] Starting to add show by Indexer Id: 457780
2026-10-06 15:34:47 DEBUG    SHOWQUEUE-ADD :: [dc17f43] 457780: Loading show info from IMDb with ID: tt39240299
2026-10-06 15:34:50 DEBUG    SHOWQUEUE-ADD :: [dc17f43] 457780: Saving to database: MAGICAL GIRL RAISING PROJECT restart
2026-10-06 15:34:53 INFO     SHOWQUEUE-ADD :: [dc17f43] Cache check completed

I confirmed persistence using read-only SQLite queries:

tv_shows:    indexer=1, indexer_id=457780, imdb_id="tt39240299"
             show_name="MAGICAL GIRL RAISING PROJECT restart"
tv_episodes: indexer=1, showid=457780, season=1, count=12

HTTP 201 establishes acceptance; the database rows and subsequent frontend-store insertion establish successful addition. The final IMDb ID follows existing external-indexer lookups; the original malformed-input regression is a separate check.

The console also contains Vue/store show is undefined errors and CSS warnings. Those frontend issues remain outside this fix. No recurrence of the original IMDb parsing exception appears in the captured add sequence.

🧪 Regression coverage and validation notes
  • I added 124 cases covering the exact failure, valid ID formats, malformed types/URLs, reassignment, conversion limits, Series fallbacks/mappings, the remaining EZTV and Trakt guards, and actual QueueItemAdd.run() completion. The 108 core parser/Series cases went from 83 failed / 25 passed against the original methods to 108 passed.
  • The full local suite passed with TZ=UTC and a process-only PATH excluding the host's ffprobe. The normal PATH exposes an existing fixture mismatch (ffprobe not available versus installed 8.0.1), also reproduced with the original methods. No test expectation was changed.
  • Compilation and git diff --check passed. Configured lint passed with the repository's global pytest exclusions; Trakt's two existing E275 warnings were confirmed against the base and excluded for its local lint check.
  • The earlier validation draft Upstream #16 completed 28 checks for dc17f43, against develop at 660bc2a. The two-commit candidate ending at 6aa24e4 passed 26 Actions jobs and both Codecov checks in draft Pull from upstream #17. I reviewed its annotations against unchanged upstream code; the existing API messages and frontend lint warnings remain unrelated to these changes.

☑️ PR checklist

  • Based on develop; separate commits for the IMDb fix and SDBits removal.
  • Read the contribution guide.
  • Verified regression tests, historical live addition, isolated runtime behavior, frontend checks, and local suites.
  • Current two-commit candidate passed fork CI; annotations reviewed and validation draft closed.

Normalize complete title IDs from paths and URLs, reject unusable values,
and guard Series fallbacks plus provider and Trakt consumers of missing IDs.

Add regression coverage for the reported TVDB value, malformed metadata,
supported formats, and completion of the actual add-queue path.
@duramato

duramato commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Just a quick note, since you've made changes to SDBits: it can be removed. The provider is down for good.

Remove its provider module, dedicated IMDb tests, frontend fixture entry,
and theme icons following review feedback on the IMDb parsing fix.

Provider discovery and legacy configuration handling need no changes.
@XxUnkn0wnxX XxUnkn0wnxX changed the title Harden IMDb ID parsing so malformed metadata cannot block adding shows Fix IMDb ID parsing and remove the discontinued SDBits provider Oct 6, 2026
@XxUnkn0wnxX

Copy link
Copy Markdown
Author

@duramato Thanks for the heads-up. I've removed SDBits in a separate commit. I included that cleanup in this PR because I'd already touched its IMDb handling.

Any other dead providers should be reviewed in a separate PR so this one stays focused on the IMDb fix and the related SDBits removal.

@XxUnkn0wnxX

Copy link
Copy Markdown
Author

Maintainers: please rerun the Windows / Node 22 frontend job once the network issue fetching dependencies from the Yarn registry is resolved, so all checks can complete successfully.

It failed during dependency installation with ESOCKETTIMEDOUT while downloading date-fns-2.30.0.tgz, before the build, lint, or tests ran. The same job passed for this exact commit in the fork validation, and the other upstream frontend jobs passed.

@medariox
medariox merged commit ecb5963 into pymedusa:develop Oct 7, 2026
37 of 38 checks passed
@medariox

medariox commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Thanks!

@XxUnkn0wnxX
XxUnkn0wnxX deleted the fix/imdb-id-parsing branch October 8, 2026 05:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants