Repository navigation
Fix --resume manifest truncation bug #2
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,65 @@ | ||
| # CLAUDE.md | ||
|
|
||
| This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. | ||
|
|
||
| ## Project Overview | ||
|
|
||
| **imgeda** is a CLI tool for exploratory data analysis (EDA) of image datasets. It scans image directories, generates JSONL manifests with metadata/pixel statistics, detects quality issues, finds duplicates, and produces visualizations. Built for both local use and AWS Lambda deployment. | ||
|
|
||
| ## Commands | ||
|
|
||
| ```bash | ||
| # Install dependencies | ||
| uv sync --all-extras | ||
|
|
||
| # Run all tests | ||
| uv run pytest | ||
|
|
||
| # Run single test file | ||
| uv run pytest tests/test_analyzer.py | ||
|
|
||
| # Run single test | ||
| uv run pytest tests/test_analyzer.py::TestAnalyzeImage::test_normal_image | ||
|
|
||
| # Run with coverage | ||
| uv run pytest --cov=src/imgeda --cov-report=html | ||
|
|
||
| # Lint | ||
| uv run ruff check src/ tests/ | ||
|
|
||
| # Lint with autofix | ||
| uv run ruff check --fix src/ tests/ | ||
|
|
||
| # Format | ||
| uv run ruff format src/ tests/ | ||
|
|
||
| # Type check (strict mode) | ||
| uv run mypy src/imgeda/ | ||
| ``` | ||
|
|
||
| ## Architecture | ||
|
|
||
| The codebase follows a layered architecture: **CLI → Pipeline → Core (pure functions)**. | ||
|
|
||
| - **`src/imgeda/cli/`** — Typer-based CLI commands (`scan`, `check`, `plot`, `report`, `info`, interactive wizard). Entry point: `cli/app.py`. | ||
| - **`src/imgeda/core/`** — Pure analysis functions with zero CLI dependencies, designed to be Lambda-compatible. Includes `analyzer.py` (single-image analysis), `detector.py` (exposure/artifact detection), `hasher.py` (perceptual hashing), `duplicates.py` (hash-based clustering with sub-hash bucketing to avoid O(n²)), and `aggregator.py` (dataset summary). | ||
| - **`src/imgeda/pipeline/`** — Orchestration layer: `ProcessPoolExecutor` parallelism with Rich progress bars, crash-tolerant resume via checkpoint logic, and graceful Ctrl+C signal handling. Batched processing with memory-bounded futures (batch size 5000). | ||
| - **`src/imgeda/io/`** — JSONL manifest I/O with atomic writes (temp file + rename) and corruption-tolerant parsing (skips malformed lines). | ||
| - **`src/imgeda/models/`** — Dataclasses with `__slots__`: `ImageRecord`, `PixelStats`, `CornerStats`, `ManifestMeta`, `ScanConfig`, `PlotConfig`. | ||
| - **`src/imgeda/plotting/`** — Eight plot types (dimensions, file_size, aspect_ratio, brightness, channels, artifacts, duplicates), each in its own module. | ||
| - **`src/imgeda/lambda_handler/`** — AWS Lambda entry point wrapping core functions. | ||
|
|
||
| ## Key Design Patterns | ||
|
|
||
| - **Core functions never raise exceptions** — corrupt/unreadable files are flagged in the `ImageRecord` rather than throwing. | ||
| - **Resume is keyed on `(path, file_size, mtime)`** — modified files are automatically re-analyzed on resume. | ||
| - **JSONL manifest format** — first line is metadata (`__manifest_meta__: true`), remaining lines are `ImageRecord` entries. Append-only with atomic metadata updates. | ||
| - **Serialization uses orjson** for performance. | ||
|
|
||
| ## Code Style | ||
|
|
||
| - Line length: 100 characters | ||
| - Target Python: 3.10+ (uses `from __future__ import annotations`) | ||
| - Type annotations throughout; mypy strict mode | ||
| - `__slots__` on all dataclasses | ||
| - Tests are class-based with pytest fixtures; test images are generated programmatically in `conftest.py` |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🛑 Logic Error: Setting
created_atto current time on resume overwrites the original scan creation timestamp. During resume, the metadata should preserve the originalcreated_atfrom the existing manifest to maintain accurate scan history. Currently, every resume resets this timestamp to "now", losing the original creation time.