alfred

Author	SHA1	Message	Date
francwa	97dc799a26	fix(tv_shows): correct PACK vs EPISODIC classification model The Phase 4 walker + rescan logic classified seasons by parser output (does the filename carry Exx?), but PACK vs EPISODIC is a structural distinction: * PACK = season folder with N flat SxxEyy videos directly inside * EPISODIC = season folder with N subfolders, each holding one video Changes: * walker.py: descends two levels under show_root and classifies each season folder by FS structure. SeasonFolder now carries mode: ReleaseMode \| None. Mixed layouts (flat + subfolders) and EPISODIC subfolders with >1 video log a warning and report mode=None. * rescan.py: trusts walker.mode; drops the bogus 'single un- numbered video → PACK with empty episodes' branch. A season with no parseable episodes is now skipped with a warning. * Tests rewritten against the real model: PACK with flat numbered files, EPISODIC with one-video-per-subfolder, malformed mixed layout skipped, single-un-numbered-file skipped. Suite: 1237 → 1245 passing.	2026-05-25 21:37:34 +02:00
francwa	fe9857aaed	docs(changelog): Phase 4 Step 5 — record dot_alfred v2 Phase 4 work Append Phase 4 entry under [Unreleased]: * Added: rescan_show v2 signature + new rescan_movie + PACK empty- episodes semantics + Settings.tmdb_cache_ttl_days + library-index anchor-mismatch warning * Removed: v1 dot_alfred stack (bridge/repository/serializer/sidecar), abstract domain ports (TVShowRepository / MovieRepository), application/library/ package, two Phase-3 quarantine test files * Internal: 1233 → 1237 passing, 10 → 8 skips; MediaWithTracks mixin parked for Phase 5 Phase 3 entries left intact (historically accurate at commit time).	2026-05-25 21:17:23 +02:00
francwa	cc334a7951	feat(dot_alfred/v2): Phase 4 Step 4 — settings + anchor warning Two small additions that close out Phase 4's loose ends. Settings — tmdb_cache_ttl_days class Settings(BaseSettings): # --- DOT_ALFRED --- tmdb_cache_ttl_days: int = 14 Default 14 days, matching the dot_alfred_v2 master spec. Will drive the Phase 5 TTL policy on TVShowLibraryIndexSidecar / MovieLibraryIndexSidecar (decide when a TMDB-cached entry is stale and triggers a refresh sync). Anchor-mismatch warning DotAlfredTVShowLibraryIndex._load_or_heal and DotAlfredMovieLibraryIndex ._load_or_heal now cross-check each indexed entry's metadata.path against the on-disk folder layout right after a successful parse. Drift (sidecar says folder X, X no longer exists under library_root) is surfaced as a WARNING log — one per missing folder, with the tmdb_id for cross-reference. No auto-heal on drift; the caller decides (the heal path remains opt-in via index.heal()). The warning fires only on the parsed-index path. The heal path always synthesizes entries from real folder names, so it can never drift — silent by construction. Tests * TestTVShowLibraryIndexAnchorWarning — 3 scenarios: warn-on-drift / no-warn-on-match / no-warn-on-heal. * TestMovieLibraryIndexAnchorWarning — symmetric coverage. Full suite: 1237 passed / 8 skipped / 4 xfailed.	2026-05-25 21:14:18 +02:00
francwa	86222d95d1	refactor(persistence): Phase 4 Step 3 — delete v1 dot_alfred + ports Now that rescan_show + rescan_movie run on the v2 release repositories (Phase 4 Steps 1-2), the v1 dot_alfred stack and its abstract domain ports have zero callers. Delete them and lift the Phase 3 quarantines. Deleted * alfred/infrastructure/persistence/dot_alfred/bridge.py * alfred/infrastructure/persistence/dot_alfred/repository.py (v1) * alfred/infrastructure/persistence/dot_alfred/serializer.py (v1) * alfred/infrastructure/persistence/dot_alfred/sidecar.py (v1) * alfred/domain/tv_shows/repositories.py (TVShowRepository ABC) * alfred/domain/movies/repositories.py (MovieRepository ABC) * tests/infrastructure/persistence/dot_alfred/test_repository.py * tests/infrastructure/persistence/dot_alfred/test_serializer.py Rewrite alfred/infrastructure/persistence/dot_alfred/__init__.py now re- exports only the v2 surface: the four concrete repositories (DotAlfredSeriesReleaseRepository, DotAlfredMovieReleaseRepository, DotAlfredTVShowLibraryIndex, DotAlfredMovieLibraryIndex) plus ShowFolderUnknown. DTO-level imports go through alfred.infrastructure.persistence.dot_alfred.v2 directly. No backwards-compat shims (per CLAUDE.md): the v1 names are gone, not aliased. Test suite drops from 10 → 8 skips (the two Phase 3 module-level skips disappear with the quarantined files). Full suite: 1233 passed / 8 skipped / 4 xfailed. The MediaWithTracks mixin in alfred.domain.shared.media is now orphaned (Episode lost its tracks in Phase 3, MovieRelease doesn't inherit it). Parked for Phase 5, which will either mount it on MovieRelease / SeasonRelease or delete it for good.	2026-05-25 21:10:32 +02:00
francwa	9e48c70b8a	feat(rescan): Phase 4 Step 2 — add rescan_movie orchestrator Mirror rescan_show for the movies library. Locates the main video via find_video_file, runs inspect_release once (movies are one-folder-one- main-file by convention), and writes a v2 MovieRelease sidecar via DotAlfredMovieReleaseRepository. Signature rescan_movie( movie_dir, , tmdb_id: TmdbId, imdb_id: ImdbId \| None = None, movie_repo: DotAlfredMovieReleaseRepository, prober, kb, ) -> MovieRelease Behavior added_at = datetime.now(UTC) — the v2 sidecar records when the release was last reconciled with disk, not filesystem mtime (which drifts across moves and hard-links). Phase 3 made this field required on MovieRelease. * No TMDB call. Index auto-heals from the new sidecar on next read. * MovieRescanFailed raised when no video is found inside movie_dir (only explicit failure mode; all other adapter errors degrade gracefully into empty / partial fields). * file_path is recorded relative to movie_dir so the sidecar stays portable across library moves. Tests tests/application/movies/test_rescan.py: 8 scenarios on the real v2 movie repo + real KB + stubbed prober. Covers track flattening, sidecar round-trip, prober returning None, video in subfolder, explicit no-video failure, imdb_id optional. Full suite: 1233 passed / 10 skipped / 4 xfailed.	2026-05-25 21:09:02 +02:00
francwa	7da0f887e7	refactor(rescan): Phase 4 Step 1 — rescan_show on v2 release repo Rewrite rescan_show to build a SeriesRelease (Phase 1 v2 aggregate) and persist it via DotAlfredSeriesReleaseRepository. The orchestrator keeps reusing inspect_release as the single source of parse/probe truth — only the assembly target changes (SeriesRelease/SeasonRelease/ EpisodeRelease instead of TVShow/Season/Episode). New signature rescan_show( show_root, , tmdb_id: TmdbId, imdb_id: ImdbId \| None = None, series_repo: DotAlfredSeriesReleaseRepository, scanner, prober, kb, ) -> SeriesRelease Identity is TMDB-anchored (tmdb_id required, no coercion); imdb_id is optional. No TMDB call from rescan — the library index auto-heals from the new sidecar on its next read. PACK vs EPISODIC Single-video + season-parsed + no-episode → SeasonRelease( mode=PACK, folder=<season folder>, episodes=()). The slot map stays empty until the Phase 5 TMDB sync supplies episode_count. We do not fabricate an EpisodeRange we cannot prove on disk. * Otherwise → EPISODIC: every file with (season, episode) becomes an EpisodeRelease with EpisodeRange(start, end) = (E, E). Multi-episode files (S01E01E02) still record only the first slot — Parser does not yet expose episode_end (existing tech debt, unchanged). Package move The orchestrator moves from alfred/application/library/ to alfred/application/tv_shows/ for symmetry with alfred/application/ movies/ (Step 2). walker.py + its tests move with it. The empty library/ package is deleted. Tests tests/application/tv_shows/test_rescan.py rewritten end-to-end on the real v2 repository, real KB, real scanner, stubbed prober. 9 happy-path + edge-case scenarios cover EPISODIC track flattening, PACK empty-episodes semantics, sidecar round-trip, imdb_id optional, empty show root, season folder with no videos, prober returning None. test_walker.py moved verbatim (import path updated). Full suite: 1214 passed / 10 skipped / 4 xfailed. The three v1 dot_alfred quarantines from Phase 3 stay in place until Step 3.	2026-05-25 21:07:25 +02:00
francwa	c22b2b78eb	refactor(domain): Phase 3 — TVShow/Movie aggregates become TMDB-only Filesystem-side concerns (file paths, tracks, quality, mode, added_at) move to the releases/ domain added in Phase 1; the TMDB aggregates now carry only identity + TMDB catalog facts. Domain entities: - TVShow: tmdb_id: TmdbId required (primary key), imdb_id: ImdbId \| None optional, status: str = "unknown" added. - Season: episode_count: int = 0 added (TMDB-cached); audio_tracks, subtitle_tracks, mode property removed. - Episode: slimmed to identity + title. file_path/file_size/tracks removed. No longer inherits MediaWithTracks. - Movie: tmdb_id required, imdb_id optional. file_path/file_size/quality/ added_at/audio_tracks/subtitle_tracks removed. get_filename() now returns "Title.Year" — quality moves to MovieRelease. Builders: - TVShowBuilder requires tmdb_id: TmdbId; imdb_id/status optional. - SeasonBuilder.set_episode_count(int) replaces set_audio_tracks / set_subtitle_tracks. No-coercion contract: TVShow(tmdb_id=1396) raises — callers pass TmdbId(1396). No ergonomic shim per the no-shims rule. Cascade fixes: - MediaOrganizer test fixtures updated to new Movie/TVShow shapes. - Movie.get_filename() re-added (without Quality) so MediaOrganizer keeps working until Phase 4 rewires it through MovieRelease. Quarantined (deleted in Phase 4 alongside v1 dot_alfred): - tests/application/library/test_rescan.py — module-level skip. - tests/infrastructure/persistence/dot_alfred/test_repository.py — module-level skip. - tests/infrastructure/persistence/dot_alfred/test_serializer.py — module-level skip. Suite: 1216 passed, 11 skipped (8 pre-existing + 3 Phase 3 quarantines), 4 xfailed. CHANGELOG updated under [Unreleased].	2026-05-25 19:54:35 +02:00
francwa	2f160644da	feat(dot_alfred/v2): bump SCHEMA_VERSION to 2 — added_at on MovieRelease Phase 3 prep: Movie aggregate is about to become TMDB-only (no filesystem fields). added_at is a release-time observation, not a TMDB-aggregate concern, so it moves to MovieRelease + MovieReleaseSidecar. - Add added_at: datetime (required) to MovieRelease with a type-check in __post_init__. - Add added_at: datetime (required) to MovieReleaseSidecar. - Bump SCHEMA_VERSION 1 → 2 with a version-history note. - Bridge round-trips added_at via Pydantic mode="json" (datetime → ISO 8601 string). - Tests: update MovieRelease fixtures, add a validator test, add an added_at round-trip test, switch hard-coded `1` assertions to SCHEMA_VERSION for future-proofing. No v1 sidecars in the wild yet — no migration code needed.	2026-05-25 19:47:25 +02:00
francwa	e65c1df229	feat(.alfred v2 — Phase 2): Pydantic sidecars, atomic repos, auto-heal index Spec: specs/dot_alfred_v2.md (Phase 2). New package alfred/infrastructure/persistence/dot_alfred/v2/: * sidecar_release.py / sidecar_root.py — Pydantic DTOs (extra="forbid", frozen=True) for per-item sidecars and the library-root index. schema_version enforced via model_validator. * serializer.py — read_yaml / atomic_write_yaml (.tmp + os.replace). SidecarSchemaError wraps YAML + Pydantic errors uniformly. * bridge.py — lossless domain <-> sidecar for SeriesRelease / MovieRelease; projection-only show_index_entry_from / movie_index_entry_from with multi-episode-file flattening. * repository.py — DotAlfredSeriesReleaseRepository / DotAlfredMovieReleaseRepository (log+skip on corruption), DotAlfredTVShowLibraryIndex / DotAlfredMovieLibraryIndex with silent auto-heal on missing/corrupt index reads. Writes never auto-heal (read paths handle that). TMDB client extensions: * TmdbSeasonInfo / TmdbShowInfo DTOs + pure parse_tv_show_info. * TMDBClient.get_tv_show_info aggregates /tv/{id} + /tv/{id}/external_ids. Domain change: * SubtitleTrack gains is_sdh: bool = False, populated from ffprobe's hearing_impaired disposition. Required for v2 sidecar parity (spec replaces v1's type: "sdh" with explicit flag). Default keeps every existing caller unchanged. Tests: 37 new v2 integration tests on tmp_path (round-trips, atomic writes, schema mismatch handling, anchor warnings, auto-heal paths) plus 16 TMDB DTO tests. Full suite: 1240 -> 1277 passed. Implementation notes filed in .claude/specs/dot_alfred_v2_notes.md (strict=True trade-off, upsert signature deviation from spec, etc.). Phases 3-5 (TVShow/Movie refactor to TMDB-only, rescan_show rewrite, v1 deletion + wiring) are next.	2026-05-25 16:01:39 +02:00
francwa	c0f6d01048	feat(releases): Phase 1 — new filesystem release domain + TmdbId VO First step of specs/dot_alfred_v2.md. Introduces a separate bounded context (alfred/domain/releases/) for the filesystem-side aggregates, disjoint from TMDB identity which stays in tv_shows/ and movies/. The link between the two worlds is TmdbId, used as the natural key in the persistence layer (no domain-level reference). New package alfred/domain/releases/: - value_objects: EpisodeRange (covers SxxE01E02E03 multi-episode files via start/end inclusive range, with count/numbers/is_single helpers), ReleaseMode enum (PACK = N video files direct in the season folder, EPISODIC = N sub-folders). - entities: TrackProfile, EpisodeRelease, SeasonRelease (with episode_count() summing each EpisodeRange.count()), SeriesRelease (tmdb_id primary anchor, optional imdb_id secondary), MovieRelease. All frozen dataclasses. - builders: SeasonReleaseBuilder + SeriesReleaseBuilder mirroring the v1 TVShowBuilder pattern. Builders sort episodes by range start on emit and reject overlapping ranges (two files claiming the same TMDB slot). from_existing() seeds a builder from an existing frozen aggregate for round-trip edits. - repositories: abstract ports (SeriesReleaseRepository, MovieReleaseRepository); concrete .alfred sidecar impls arrive in Phase 2. New shared VO alfred/domain/shared/value_objects.py::TmdbId — positive int, rejects bool/str/float, symmetric with the existing ImdbId VO. 73 unit tests cover VO validation, entity invariants, builder sort + overlap detection, and from_existing() round-trips. v1 code paths are untouched at this stage; the new domain coexists with the old TVShow aggregate until Phase 3 refactors it.	2026-05-25 15:19:23 +02:00
francwa	de7030fa9c	feat(library): add rescan_show orchestrator + walker (Step 4) Step 4 of specs/dot_alfred.md — rebuild a TVShow aggregate from disk by reusing the existing release pipeline (inspect_release) on every video file in a show folder, then persist via the .alfred repository. - alfred/application/library/walker.py — pure structural walk (season folders detected via \bS\d{1,2}\b regex, video files filtered against kb.video_extensions, no recursion). - alfred/application/library/rescan.py — orchestrator that ingests each season folder, infers PACK vs EPISODIC from on-disk file count + parser output, and assembles via TVShowBuilder. Episode paths stored relative to show_root. Logs + skips corrupt input (no season parsed, mixed season numbers, unparseable episodes). - Season now inherits MediaWithTracks: PACK seasons carry season-level audio_tracks / subtitle_tracks; EPISODIC seasons leave them empty (tracks live per-episode). SeasonBuilder gains set_audio_tracks / set_subtitle_tracks; bridge writes/reads them in the PACK branch via shared _synth_* helpers. Out of scope, tracked as tech debt: adjacent .srt capture, multi- episode (episode_end), TMDB-driven PACK detection (the current heuristic '1 file == PACK' is a placeholder until ShowTracker lands). 18 new tests (11 walker + 7 rescan integration) on tmp_path with the Foundation layout. Full suite: 1149 passed.	2026-05-24 15:22:18 +02:00
francwa	3622c95154	chore(lint): Lint the shit out of it	2026-05-24 15:21:58 +02:00
francwa	c7c11180d9	feat(persistence): add DotAlfredTVShowRepository (filesystem-backed) Step 3 of specs/dot_alfred.md. Concrete TVShowRepository implementation reading and writing per-show .alfred YAML files under a configurable library_root. Writes are atomic (.alfred.tmp + os.replace), reads tolerate corrupted/wrong-schema sidecars (log + skip), and the repo never invents a folder name — save(show) requires the target folder to exist beforehand (raises ShowFolderUnknown otherwise), matching the spec's MediaOrganizer-then-sidecar split. Cold folders without a sidecar are skipped by find_all and yield None from find_by_imdb_id — the upcoming rescan_show tool (step 4) will own the opt-in rebuild path. A small bridge module translates between the rich domain TVShow (AudioTrack/SubtitleTrack with full ffprobe minutiae) and the compact sidecar shape (language-only audio, embedded-only subs with type derived from is_forced). The bridge is intentionally lossy on probe details the sidecar does not store, per the spec's factual-only philosophy. 20 integration tests on tmp_path: round-trip save/find, cold-folder/unknown-id returns, find_all skipping (corrupted/schema-violating sidecars), delete/exists, atomic write (no .alfred.tmp leftover), overwrite, and folder-name fallbacks (get_folder_name guess + full-scan rescue when renamed).	2026-05-22 17:16:41 +02:00
francwa	b0e275bd11	feat(persistence): add .alfred sidecar serializer (DTO ↔ dict) Step 2 of the specs/dot_alfred.md plan. Pure-dict in/out (serialize(sidecar) -> dict, deserialize(data) -> ShowSidecar); YAML I/O lives in the repository layer (step 3) and is kept out for trivial testability. DTOs mirror the YAML schema field-for-field: - ShowSidecar (root: imdb_id, tmdb_id, schema_version, seasons) - SeasonSidecar (number, path, optional audio/subtitles, optional episodes) - EpisodeSidecar (number, path, optional audio/subtitles) - SubtitleEntry (language, source, type) The sidecar acts as a scan cache: it stores only what is genuinely costly to recompute — folder/file paths (skipping the FS walk) and probed track metadata (skipping ffprobe). Release identifiers (group, source, quality, codec) live in folder/file names and are derived on demand by the parser; they are deliberately absent from the schema and rejected as unknown keys on deserialize. The serializer is strict on schema: unknown keys at any level raise SidecarSchemaError, missing required fields raise clearly, and bool cannot sneak in as a season/episode number. Optional fields (tmdb_id, empty audio/subtitles/episodes) are omitted from the output rather than emitted as null / []. Tests cover round-trip equivalence (DTO → dict → DTO and DTO → YAML text → DTO), the Foundation S01 PACK case (real-world fixture with mixed sub types — superset captured at season scope), and a Breaking Bad S05 EPISODIC case. An on-disk tmp_path fixture recreates the Foundation folder structure with placeholder files, ready to be reused by the upcoming repository walk tests in step 3.	2026-05-22 16:56:56 +02:00
francwa	6c12c18a27	refactor(tv_shows): freeze aggregate, builder-only construction, drop ShowTracker fields The TVShow aggregate is now fully immutable. TVShow, Season and Episode are @dataclass(frozen=True), children stored as ordered tuples sorted by number. All construction goes through TVShowBuilder / SeasonBuilder (new module), which expose from_existing() to seed from a current frozen aggregate and apply modifications. ShowTracker-territory fields are stripped from the domain: ShowStatus, CollectionStatus, expected_seasons/episodes, aired_episodes, collection_status(), is_complete_series(), missing_episodes(), is_ongoing(), is_ended(), Season.name, the aired<=expected validation, and the TMDB status string mapping. These will reappear in a dedicated ShowTracker layer (to be designed) combining the .alfred sidecar with live TMDB data. New SeasonMode enum (PACK / EPISODIC) computed at read time from the season's structural shape — never stored, the YAML sidecar encodes the mode via presence/absence of the episodes: block. Test suite for the domain entirely rewritten to cover frozen invariants, builder ordering, last-write-wins, from_existing round-trip, and SeasonMode derivation. Full suite still green (1078 passed).	2026-05-22 16:09:37 +02:00
francwa	1427c8a54b	docs(specs): add dot_alfred sidecar design doc First entry in the new specs/ directory. Specifies the layout and semantics of the per-show .alfred/ sidecar that will back the future concrete TVShowRepository: - One .alfred/ directory per show, containing show.yaml + one season_NN.yaml per season (zero-padded, season_00 for Specials). - Per-episode entries store file size + mtime so cache lookups skip a full ffprobe rescan when nothing changed. - Self-healing on drift (file missing/modified/new) without raising. - Atomic writes via temp file + os.replace(). - Phased implementation plan (builder + freeze first, then serializer, then cache validator, then repo, then wiring). No code yet — spec only, awaiting review before the implementation phases. Companion entry in CHANGELOG (Added).	2026-05-21 18:05:55 +02:00
francwa	8491edac22	infra(gitignore): track specs/ + carve out private .claude/ The repo-level .gitignore had a blanket *.md rule with only CHANGELOG.md exempted. Two adjustments: - Allow specs/ to be tracked (design docs / RFCs live here, public). - Restrict the README.md exception to the root (/README.md) so that per-directory README files (e.g. tests/fixtures/releases/README.md) stay ignored as before — no unintended scope creep. - Explicitly ignore /.claude/, the private dev-docs sub-repo that lives inside the working tree but is versioned and pushed separately. CHANGELOG: Internal entry.	2026-05-21 18:05:33 +02:00
francwa	02e478a157	refactor(domain): freeze Movie and Episode, switch track collections to tuple Movie and Episode become @dataclass(frozen=True, eq=False), with audio_tracks/subtitle_tracks held as tuple[...] instead of list[...]. Identity-based equality is preserved via the existing __eq__/__hash__. __post_init__ coercion (imdb_id, title, season_number, episode_number) uses object.__setattr__ to stay compatible with frozen. The MediaWithTracks mixin contract is updated to tuple accordingly. Callers projecting enrichment results (probe output, file metadata) now rebuild via dataclasses.replace(...) — same pattern recently adopted for ParsedRelease. Season and TVShow stay mutable for now: freezing the aggregate root would cascade a full reconstruction on every add_episode, deferred.	2026-05-21 13:40:22 +02:00
francwa	3dc73a5214	feat(release): add fullwidth vertical bar ｜ (U+FF5C) to separators CJK release names sometimes use the fullwidth vertical bar as a token separator, as do occasional decorative YouTube-style uploads. Adding the codepoint to separators.yaml lets the tokenizer split on it instead of leaving the wide pipe glued onto an adjacent token. The tokenizer in alfred/domain/release/parser/pipeline.py iterates the separator list as plain strings (no regex), so a multi-byte UTF-8 separator works without any code change.	2026-05-21 08:05:56 +02:00
francwa	88f156b7a4	refactor(subtitles): rename SubtitleCandidate → SubtitleScanResult The old name conflated 'might become a placed subtitle' with 'what a scan pass produced'. The class is the output of a scan/identify pass — language/format may still be None while classification is in progress, confidence reflects classifier certainty, raw_tokens holds filename fragments under analysis. SubtitleScanResult says that directly. Pure rename + refreshed docstring; no behavior change. Touches the domain entity, the matcher/identifier/utils services, the manage_subtitles use case, the placer, the metadata store, the shared-media cross-ref comment, and 7 test modules.	2026-05-21 08:05:46 +02:00
francwa	5107cb32c0	feat(release): InspectedResult.recommended_action centralizes exclusion decision Add a derived 'recommended_action' property on InspectedResult that collapses the orchestrator's go / wait / skip decision into one value: - 'skip' → no main_video, or media_type == 'other' - 'ask_user' → media_type == 'unknown', or road == 'path_of_pain' - 'process' → confident parse with a main video on disk The ordering is part of the contract (skip > ask_user > process) — documented in the property docstring. Until now every consumer (workflows, the agent, the orchestrator sketch) had to re-derive this from the road / media_type / main_video triple, with subtle drift between sites. One place, one rule. Exposed through the analyze_release tool so the LLM can route on it. Spec YAML updated to describe the new field. Suite: 1083 passed (+6 new tests in tests/application/test_inspect.py covering the four branches and the precedence rules).	2026-05-21 07:54:17 +02:00
francwa	b7979c0f8b	refactor(release): freeze ParsedRelease + enrich_from_probe returns new instance ParsedRelease is now @dataclass(frozen=True). The enrichment passes that used to patch fields in place now produce new instances: - enrich_from_probe(parsed, info, kb) returns a new ParsedRelease via dataclasses.replace (no allocation when no field changed). - inspect_release rebinds 'parsed' after detect_media_type (wrapped in MediaTypeToken — the strict isinstance check now also runs on replace) and after enrich_from_probe. languages becomes a tuple[str, ...] so the VO is properly immutable. Parser pipeline packs languages as a tuple in the assemble dict. Callers updated: inspect_release, testing/recognize_folders_in_downloads.py. Tests updated: 22 enrich_from_probe call sites rebound, language assertions switched to tuple literals, test_release_fixtures normalizes result['languages'] back to list for YAML-fixture comparison. Suite: 1077 passed.	2026-05-21 07:51:49 +02:00
francwa	9f1ce94690	refactor(application): inject kb/prober into resolve_destination use cases Remove the module-level _KB / _PROBER singletons from alfred/application/filesystem/resolve_destination.py. The four resolve_{season,episode,movie,series}_destination use cases now take kb: ReleaseKnowledge and prober: MediaProber as required arguments, matching the shape of inspect_release. The singletons now live at the agent-tools frontier (alfred/agent/tools/filesystem.py), where the LLM-facing wrappers instantiate YamlReleaseKnowledge / FfprobeMediaProber once and thread them through. The wrappers' Python signatures are unchanged — the inspect-based JSON-schema generator in agent/registry.py still sees the same LLM-passable params. analyze_release drops the dirty 'from ... import _KB' indirection. Tests inject their own stubs by keyword (prober=_StubProber(...)) via thin convenience wrappers, replacing the prior monkeypatch.setattr(rd, '_PROBER', ...) pattern. testing/debug_release.py: instantiate YamlReleaseKnowledge() / FfprobeMediaProber() inline at the two call sites. Suite: 1077 passed.	2026-05-21 07:46:13 +02:00
francwa	5e0ed11672	refactor(release): rename ParsePath enum to TokenizationRoute ParsePath collided with pathlib.Path in mental models, and was one letter from the parse_path attribute that stores its value — confusion on confusion. Road (EASY/SHITTY/PATH_OF_PAIN) is the parser-confidence axis; TokenizationRoute (DIRECT/SANITIZED/AI) is the tokenization-method axis. They're orthogonal and the new name makes that obvious. Field name parse_path stays — it's the right name for the attribute that holds the route. String values ("direct", "sanitized", "ai") stay too, so YAML fixtures and the analyze_release tool spec are unchanged. Only the type symbol changes: - value_objects.py: class rename + docstring spelling out orthogonality with Road. - services.py: 3 call sites. - scoring.py: docstring cross-reference updated. - tests/domain/release/test_parser_v2_scoring.py: import + 3 call sites.	2026-05-21 07:39:42 +02:00
francwa	0246f85ef8	refactor(release): move codec mappings from code to YAML knowledge The three module-level dicts in enrich_from_probe (ffprobe codec name to scene token, channel count to layout) were exactly the kind of domain lookup table CLAUDE.md says belongs in YAML, not in Python. Move them to alfred/knowledge/release/probe_mappings.yaml, load through a new ReleaseKnowledge.probe_mappings port field, and add a kb parameter to enrich_from_probe so the consumer reads the maps via the same injection pattern as everything else. - New knowledge file: alfred/knowledge/release/probe_mappings.yaml - New loader: load_probe_mappings() in infrastructure/knowledge/release.py (normalizes channel-count keys back to int). - Port: ReleaseKnowledge gains probe_mappings: dict. - Adapter: YamlReleaseKnowledge populates it at __init__. - Consumer: enrich_from_probe(parsed, info, kb) reads the three sub-maps from kb.probe_mappings; unknown codecs still fall back to uppercase raw value, same behaviour as before. - Call sites updated: inspect_release passes kb through; the testing script gets its kb wiring (it was already broken since the ReleaseKnowledge refactor); all 22 enrich_from_probe call sites in tests/application/test_enrich_from_probe.py pass _KB.	2026-05-21 07:37:42 +02:00
francwa	e62dc90bd1	refactor(release): make tech_string a derived property ParsedRelease.tech_string was a stored str field re-computed in two places (assemble() at parse time, enrich_from_probe() after the probe). The second site was a reactive fix (`e79ca46`) for filename builders that saw a stale value. Turn it into an @property so it stays in sync with quality/source/codec by construction. - Drop the field from the dataclass + the key from assemble()'s dict. - Drop tech_string="" from parse_release's malformed-name fallback. - Drop the manual recomputation at the end of enrich_from_probe. - Inject the property into asdict() result in the fixtures runner (same treatment as is_season_pack). - Update tests that passed tech_string= to the constructor; rewrite the TestTechString case that mutated p.tech_string manually.	2026-05-21 07:33:53 +02:00
francwa	688c37bbec	docs(changelog): recap session 2026-05-20 tech-debt cleanup Consolidate the five domain-purity refactors of the session under [Unreleased]: RuleScopeLevel enum, FilePath VO post_init, Language strict + from_raw, ParsedRelease.normalised → clean, ParsedRelease enum strictness. Removes the duplicate min_movie_size_bytes entry (now sits under its proper Removed section).	2026-05-20 23:57:06 +02:00
francwa	757e4045ee	refactor(release): ParsedRelease.media_type & parse_path are strict enums The fields were already typed as MediaTypeToken / ParsePath, but a tolerant __post_init__ coerced raw strings into their enum form. With MediaTypeToken(str, Enum) (and ParsePath idem), the coercion served no purpose — callers that pass '.value' got back the enum anyway, and callers that pass an unknown string got a ValidationError just like they would now. Strict mode: constructor rejects non-enum values directly. The two in-tree builders (parse_release() and the parser pipeline) already produce enum values; all .value sites have been removed. Drops the unused _VALID_MEDIA_TYPES / _VALID_PARSE_PATHS lookup tables.	2026-05-20 23:52:30 +02:00
francwa	c3767aacb6	refactor(release): rename ParsedRelease.normalised → clean Le champ s'appelait normalised mais ne faisait pas la normalisation suggérée par son nom (dots instead of spaces). En pratique il contient raw - site_tag - apostrophes, qui sert uniquement à season_folder_name() via _strip_episode_from_normalized. Renommé en 'clean' qui décrit ce qu'il contient réellement, docstring corrigée.	2026-05-20 23:50:05 +02:00
francwa	5bcf22b408	refactor(shared): Language VO is strict; from_raw() factory for un-normalized input object.__setattr__ inside __post_init__ on a frozen dataclass is a code smell — it bypasses the immutability guarantee to mutate fields mid-construction. Split the responsibilities: * Direct constructor is strict — rejects un-normalized input (uppercase iso, whitespace in aliases, etc.) so once a Language exists in the system, its fields are guaranteed canonical. * Language.from_raw() factory handles arbitrary YAML/user input — it lowercases the iso, dedups/normalizes aliases, then constructs. Only caller that built from raw data (LanguageRegistry loading YAML) moves to from_raw(). Test fixtures already pass normalized data so they keep using the direct constructor.	2026-05-20 23:48:30 +02:00
francwa	cfa9f54d9f	refactor(shared): FilePath VO uses __post_init__ instead of custom __init__ Custom __init__ on a @dataclass(frozen=True) is a code smell — it bypasses the generated dataclass __init__ and re-implements the str/Path coercion + frozen-aware setattr by hand. Replaced with a single __post_init__ that performs the same normalization. Same public API (FilePath(str) and FilePath(Path) both work), same behavior, no callers touched.	2026-05-20 23:47:03 +02:00
francwa	f0aaf50c97	refactor(subtitles): RuleScope.level → RuleScopeLevel enum Six niveaux possibles (global, release_group, movie, show, season, episode) étaient passés en str libre, le commentaire docstring servant de seule documentation. Introduit RuleScopeLevel(str, Enum) — toujours sérialisable en YAML, mais le set fixe est désormais imposé par le typage. to_dict() sort explicitement .value pour rester safe côté écrivains YAML.	2026-05-20 23:46:22 +02:00
francwa	a09262b33f	chore(settings): remove unused min_movie_size_bytes Le champ + son validator étaient orphelins depuis la suppression de MovieService.validate_movie_file. L'exclusion par extension (application/release/supported_media.py) + le PoP couvrent désormais la règle 'vrai film vs sample'. Si on a un jour besoin d'un seuil de taille, il ira dans data/knowledge/, pas dans settings.	2026-05-20 23:41:41 +02:00
francwa	9c7cd66d2b	Merge branch 'refactor/flatten-shared-media'	2026-05-20 23:35:52 +02:00
francwa	83dbed887b	refactor(domain): flatten shared/media package into single module Six small files (audio, video, subtitle, info, matching, tracks_mixin + __init__) collapsed into one ~250 LoC media.py module. Python treats media.py and media/__init__.py interchangeably, so the 12 import sites that read 'from alfred.domain.shared.media import ...' continue to work without changes. Reasoning: the whole bounded context fits on one screen; splitting into sub-modules added more navigation friction than it saved. Tests stay green (1077 passed).	2026-05-20 23:35:49 +02:00
francwa	0c9489e16b	Merge branch 'feat/parser-phase-d'	2026-05-20 23:30:36 +02:00
francwa	621bb96995	fix(release/parser): pre-strip apostrophes so titles like Don't parse cleanly Apostrophes are in the forbidden-chars list, which made any release with a title like "Don't" or "L'avare" short-circuit to the AI fallback (parse_path=ai, everything UNKNOWN). They are now stripped up front from the name before the well-formed check and tokenize, so the parse completes normally. The raw name is preserved on the VO; only the title field loses its apostrophe. parse_path becomes 'sanitized' when an apostrophe was stripped, to surface that the parser cleaned something up. Fixtures updated: - shitty/honey_uhd_hdr/ — went from total UNKNOWN to a clean parse (title=Honey.Dont, year=2025, quality=2160p, source=WEBRip, codec=x265, group=Amen). - path_of_pain/the_prodigy_full_chaos/ — went from total failure to partial success (title, year, source, codec extracted). Remaining gaps (1080i, multi-word audio, Blu-ray-with-dash) are tracked separately in tech debt.	2026-05-20 23:29:10 +02:00
francwa	448ef3b79c	fix(release/parser): recognize Sxx-yy season range as tv_complete `Der.Tatortreiniger.S01-06.GERMAN...` previously parsed as a movie with 'S01-06' glued to the title. The parser now matches the season-range form in _parse_season_episode (returning season=first, episode=None), and the assemble step detects the range token to promote media_type to 'tv_complete'. The first season is exposed as `season` so `is_season_pack` fires (season is not None and episode is None) — useful for routing to a series root folder. Fixture shitty/tatortreiniger_flat_multiseason/ updated: - title: Der.Tatortreiniger.S01-06 → Der.Tatortreiniger - season: null → 1 - media_type: movie → tv_complete - is_season_pack: false → true	2026-05-20 23:26:40 +02:00
francwa	b1c7f35ffb	fix(release/parser): drop pure-punctuation TITLE tokens at assembly Releases using ' - ' as a separator (Vinyl - 1x01 - FHD) tokenize to ['Vinyl', '-', '1x01', '-', 'FHD'] — the standalone '-' tokens were ending up in title_parts and leaked into the joined title ('Vinyl.-'). We can't add '-' to the separator list (it would break codec-GROUP), so we filter at assembly: a TITLE token with no alphanumeric characters carries no title content. Side win: same logic eliminates the UTF-8 wide-pipe '｜' from the khruangbin_yt_wide_pipe fixture title. Fixtures updated: - shitty/vinyl_1x01_format/expected.yaml (title: Vinyl.- → Vinyl) - path_of_pain/khruangbin_yt_wide_pipe/expected.yaml (｜ dropped)	2026-05-20 23:24:40 +02:00
francwa	5bbdc9081f	fix(release/parser): collapse chained multi-episode markers to full range S14E09E10E11 previously parsed to episode=9, episode_end=10 — E11 was silently dropped. The parser now takes episodes[-1] as episode_end so the full chain is captured (episode=9, episode_end=11). Intermediate values stay implied. Fixture shitty/archer_multi_episode/ updated from anti-regression of the bug to anti-regression of the fix.	2026-05-20 23:23:08 +02:00
francwa	5d7b214af2	Merge branch 'refactor/language-port'	2026-05-20 23:20:18 +02:00
francwa	18267d0165	refactor(language): LanguageRepository port + SubtitleKnowledgeBase wired to it Mirror the MediaProber / FilesystemScanner pattern for language lookup: - New Protocol `LanguageRepository` in alfred.domain.shared.ports covering from_iso, from_any, all, __contains__, __len__ — the surface previously coupled to the concrete LanguageRegistry. - SubtitleKnowledgeBase types its `language_registry` parameter against the Protocol; the concrete LanguageRegistry stays in infrastructure as the YAML-backed adapter and remains the default when no repository is injected. - New unit tests in tests/infrastructure/test_language_registry.py cover the adapter surface (from_iso, from_any, membership, case-insensitivity, non-string inputs). Behaviour is unchanged for existing callers. The split opens the door to in-memory fakes in future tests without loading the full ISO 639 YAML.	2026-05-20 23:18:25 +02:00
francwa	19fe8a519a	Merge branch 'feat/release-inspect-orchestrator' Inspection pipeline groundwork: - MediaProber.probe() port extension (full media inspection on the port) - inspect_release orchestrator + InspectedResult frozen VO - enrich_from_probe now refreshes tech_string - resolve_*_destination use cases consume inspect_release - detect_media_type & enrich_from_probe moved to application/release	2026-05-20 09:31:22 +02:00
francwa	a0d1846ff2	refactor(release): move detect_media_type & enrich_from_probe to application/release Both helpers are inspection-pipeline pieces, not filesystem use cases — they belong next to inspect_release, not next to move_media / resolve_destination / list_folder. The move also kills the lazy import that was hiding inside _resolve_parsed: alfred.application.filesystem.resolve_destination no longer triggers a cycle through alfred.application.filesystem __init__ when loading inspect_release. Top-level import restored. Call sites updated: inspect.py, test_detect_media_type.py, test_enrich_from_probe.py, testing/recognize_folders_in_downloads.py. Module docstrings + test-file docstrings updated to match the new location.	2026-05-20 09:29:58 +02:00
francwa	0fb59a4581	feat(filesystem): wire inspect_release into resolve_destination The four resolve__destination use cases now route through a private _resolve_parsed helper that picks the right entry point: - source path provided AND it exists -> inspect_release(name, path) runs the full pipeline (parse + media-type refinement + probe + enrich), so missing tech tokens (quality, codec, ...) get filled by ffprobe and the refreshed tech_string lands in the destination folder / file names. - source path missing or absent -> parse_release(name) only, same behavior as before. Back-compat: tests using fake /dl/.mkv paths still pass unchanged. resolve_episode_destination / resolve_movie_destination reuse their existing source_file parameter as the inspection target. The two folder-move use cases (season / series) gain a new OPTIONAL source_path parameter — threaded through the agent tool wrappers and documented in the YAML specs. The lazy import inside _resolve_parsed avoids a circular import: inspect_release imports detect_media_type / enrich_from_probe from the same application.filesystem package whose __init__ re-exports resolve_destination. Three new tests in TestProbeEnrichmentWiring with a stub MediaProber prove the wiring: movie picks up probe quality, season picks it up via source_path, and a missing path correctly skips probe (back-compat guard).	2026-05-20 09:26:30 +02:00
francwa	e79ca462b8	fix(release): refresh tech_string after enrich_from_probe enrich_from_probe fills None fields on ParsedRelease (quality, source, codec, audio_*, languages) but left tech_string at its parser-time value — so the filename builders (movie_folder_name, episode_filename, …) saw stale tech tokens even after a successful probe. Re-derive tech_string the same way the parser does — quality.source.codec joined by dots, skipping None — at the end of enrich_from_probe. Token- level values still win because enrich only fills None fields. Four new tests in TestTechString cover: enrichment rebuilds it, existing source survives, no-info input leaves it untouched, fully empty parsed produces ''.	2026-05-20 09:26:09 +02:00
francwa	03aa844d7d	feat(release): inspect_release orchestrator + InspectedResult VO New application-layer entry point that composes the four inspection layers in one call: 1. parse_release(name, kb) -> (ParsedRelease, ParseReport) 2. detect_media_type(parsed, path, kb) -> patch parsed.media_type 3. find_main_video(path, kb) -> Path \| None (top-level scan) 4. prober.probe(video) + enrich -> when video exists and media_type not in {unknown, other} Returns a frozen InspectedResult(parsed, report, source_path, main_video, media_info, probe_used). kb and prober are injected — no module-level singletons in inspect.py. analyze_release tool now delegates to inspect_release; its output gains two fields, confidence (0-100) and road (easy/shitty/path_of_pain), surfaced from ParseReport so the LLM can route by confidence. Spec updated to document them. 12 new tests covering happy paths, probe gating (no video, media_type 'other', probe failure), mutation contract (detect refining parsed.media_type, enrich filling None fields), resilience (nonexistent path), and frozen contract. Suite: 1058 passing.	2026-05-20 09:15:29 +02:00
francwa	c303efea48	refactor(probe): consolidate full probe() into MediaProber port Add probe(video) -> MediaInfo \| None to the MediaProber Protocol and implement it on FfprobeMediaProber. The standalone alfred/infrastructure/filesystem/ffprobe.py module is removed; all callers (analyze_release / probe_media tools, testing scripts) now go through the adapter. Tests for the probe path moved to tests/infrastructure/test_ffprobe_prober.py (patching subprocess.run at the adapter module level). Unblocks the upcoming inspect_release orchestrator, which needs the port — not a free function — to compose parse + main-video selection + probe in one shot.	2026-05-20 09:11:24 +02:00
francwa	5db350a1df	Merge branch 'feat/release-parser-scoring'	2026-05-20 08:47:38 +02:00
francwa	12dc796ea2	docs(changelog): freeze confidence scoring + exclusion work block	2026-05-20 08:47:29 +02:00

1 2 3 4 5

213 Commits