diff --git a/.cursor/plans/ewfinfo_parity_port_ff7d3303.plan.md b/.cursor/plans/ewfinfo_parity_port_ff7d3303.plan.md new file mode 100644 index 0000000..e0b4d82 --- /dev/null +++ b/.cursor/plans/ewfinfo_parity_port_ff7d3303.plan.md @@ -0,0 +1,140 @@ +--- +name: ewfinfo parity port +overview: Implement a Rust `ewfinfo` CLI (clap) + supporting library report/printer APIs in `crates/ewf` that match libewf’s `ewfinfo` image-metadata behavior/output (text + DFXML), with explicit TODO/`unimplemented` for any unsupported surface area (no silent fallbacks). Keep logical file outputs (`-F`/`-H`/`-B`) in the `ewfinfo` binary target (not the library); use `miette` for application-facing diagnostics while keeping library errors in `thiserror`. +todos: + - id: ewfinfo-api + content: Add documented `crates/ewf::ewfinfo` library module for image-metadata reports + printers (no LEF/-F/-H/-B). + status: completed + - id: docs-and-unit-tests + content: Add module docs + rustdoc examples + unit tests for every new public API (include “References” with upstream source file paths). + status: completed + - id: ewf1-metadata + content: Implement EWF1 metadata extraction for header values, media/ewf info, digest hashes, sessions/tracks, and acquisition errors. + status: completed + - id: ewf2-metadata + content: Implement EWF2 metadata extraction (device/case tags, set-id, compression, md5/sha1 sections, etc.). + status: completed + - id: ewfinfo-logical-cli + content: Implement `ewfinfo` CLI-only logical evidence outputs (`-F`/`-H`/`-B`) in the `ewfinfo` binary target (may extend `LefReader` public API minimally, but keep formatting + bodyfile semantics out of the library). + status: completed + - id: printers + content: Implement text + DFXML printers for the image-metadata report that match libewf `ewfinfo` formatting. + status: completed + - id: cli-ewfinfo + content: Add `ewfinfo` binary target using clap for libewf-compatible flags/conflicts and miette for user-facing errors (binary can be multi-file). + status: completed + - id: golden-tests + content: Add golden-output tests for text/dfxml/file-entry/hierarchy/bodyfile, plus TODO/unimplemented tests for unsupported paths. + status: completed +--- + +# Port libewf `ewfinfo` to `crates/ewf` + +## Goal + +- Add a **new `ewfinfo` Rust binary** (in `crates/ewf`) and the supporting **public library APIs** so we can reproduce libewf `ewfinfo` feature-for-feature: +- Options: `-A -B -d -e -f -F -H -i -m -s -v -V -h` (per `external/libewf/manuals/ewfinfo.1` + `external/libewf/ewftools/ewfinfo.c`). +- Output: **text** and **DFXML** with the same section structure + formatting. +- **No best-effort fallbacks**: if something isn’t implemented in Rust, we leave an explicit `TODO:` and return `unimplemented!()` / `Error::Unsupported("TODO: ...")` rather than silently degrading. + +## Approach (map C ewfinfo → Rust) + +### 1) Create a reusable ewfinfo library module (image metadata only) + +- Add a new **library** module under `crates/ewf/src/ewfinfo/` that provides a Rust-native “report + printer” API for **image metadata only**. +- **Do not** 1:1 port or “mirror” libewf’s `info_handle_t`. The `ewfinfo` **binary target** should own the clap-facing types and translate them into a small, strongly-typed library API. +- Keep the boundary sharp: +- **Library (`crates/ewf`)**: build a structured report for EWF *image* metadata + print it (text/DFXML). +- **Binary target (`ewfinfo`)**: owns **logical evidence** modes and outputs (`-F`/`-H`/`-B`), path separator handling, and any bodyfile semantics. +- Proposed (public) library surface (names TBD; document every `pub` item): +- `EwfInfoReport`: data model for the sections libewf prints for images: +- Acquisition/header values (libewf title: “Acquiry information”) +- EWF information +- Media information +- Digest hash information +- Sessions / Tracks +- Acquisition read errors +- `EwfInfoPrinter` (trait) + concrete printers (e.g. `TextPrinter`, `DfxmlPrinter`) with `EwfInfoPrintOptions` (date formatting, verbosity, etc.) +- `EwfInfoBuildOptions` for report construction knobs that actually affect parsing/normalization (e.g. header decoding/codepage), **not** CLI-only options like `-s`/`-B`. +- `EwfInfoError` (library) implemented with `thiserror`. +- **Module documentation requirements** (non-negotiable): +- Each new module gets `//!` docs with a short compatibility statement and a “References” section that attributes upstream reference material by file path (at minimum): +- `external/libewf/ewftools/info_handle.h` +- `external/libewf/ewftools/ewfinfo.c` +- `external/libewf/manuals/ewfinfo.1` +- Include rustdoc examples (doctests) that exercise the public API surface (using existing small fixtures/builders). + +### 2) Extend readers to expose the metadata ewfinfo prints + +Keep the existing small summary API (`EwfInfo` in [`crates/ewf/src/info.rs`](crates/ewf/src/info.rs)) stable; add *new* APIs instead. + +#### Disk images (`EwfReader`) + +- Add `EwfReader::ewfinfo_report(&self, opts: &EwfInfoBuildOptions) -> Result`. +- Implement format-specific extraction: +- **EWF1 (E01/S01)**: parse required sections from the already-discovered section descriptors (header/header2/volume/disk/data/hash/digest/error/session/track). +- Header values: parse both `header` (ASCII/codepage) and `header2` (UTF-16LE) and construct the same identifier→description mapping used by `info_handle_header_values_fprint`. +- EWF + media info fields: derive from parsed volume/disk/data structures (`sectors_per_chunk`, `bytes_per_sector`, `number_of_sectors`, `error_granularity`, `set_identifier`, compression level/method). +- Hash values: read stored global hashes from digest/hash sections (no recomputation unless ewfinfo does so). +- Sessions/tracks: parse ranges as start_sector/sector_count. +- Acquisition errors: parse ranges as start_sector/sector_count. +- **EWF2 (Ex01)**: reuse existing parsing in [`crates/ewf/src/reader.rs`](crates/ewf/src/reader.rs) (case data/device information tags) to populate the same report fields: +- `set_id`, `compression_method`, `chunk_count`, `sectors_per_chunk`, `bytes_per_sector`, `number_of_sectors` +- global MD5/SHA1 sections (parse from section types) to populate digest hash info +- sessions/tracks/errors: if format doesn’t carry them, report 0 entries (matching libewf behavior for “none present”). + +### 3) Implement printers for exact text + DFXML output + +- Add printer modules under `crates/ewf/src/ewfinfo/`: +- Text printer replicating: +- section headers/footers and indentation +- field label padding (the C code aligns to 24 columns) +- exact section titles: “Acquiry information”, “EWF information”, “Media information”, “Digest hash information”, etc. +- DFXML printer replicating the XML emitted by `info_handle_dfxml_*_fprint` (header/footer + element names). +- **No fallback behavior**: invalid inputs/options should be rejected early. For the **CLI**, clap should enforce as much as possible (enums, conflicts, defaults). For the **library**, return explicit `EwfInfoError::Unsupported("TODO: …")` where needed rather than silently defaulting. + +### 4) Add the `ewfinfo` binary target (clap + miette) and keep logical outputs there + +- Implement `ewfinfo` as a **binary target** that can be split across multiple Rust modules (prefer directory-style bin: `crates/ewf/src/bin/ewfinfo/main.rs` + submodules). +- Use **clap** to translate libewf flags/idioms into a typed CLI surface (instead of manually porting structs): +- `#[derive(Parser)] `root + `Args`/`Subcommand` as needed. +- `ValueEnum` / typed enums for `-f` (text/dfxml), `-d` (date format), etc. +- conflict groups for `-e`/`-i`/`-m` (mutually exclusive), and for logical modes (`-F` vs `-H` etc.) as required. +- rely on clap’s generated `--help`/`--version` UX while keeping short flags compatible. +- Use **miette** for user-facing diagnostics: +- Map library `thiserror` errors into `miette::Diagnostic` at the application boundary with helpful context (`wrap_err`, filenames, option values). +- Keep **logical evidence outputs** out of the library: +- `-F` (file entry detail), `-H` (hierarchy), `-B` (bodyfile) live in the `ewfinfo` binary target. +- If the binary needs additional LEF accessors, add small, generic `pub` APIs to `LefReader` (document + unit test them), but keep formatting and bodyfile semantics in the binary. + +### 5) Tests + documentation (unit tests first, then golden outputs) + +- Add **unit tests** for every new library type/module under `crates/ewf/src/ewfinfo/` (and any new public reader accessors): +- parsing/normalization invariants +- section ordering and required fields presence +- printer formatting invariants (labels, indentation, titles) +- Add **CLI unit tests** (clap `try_parse_from`) for flag conflicts/defaults and for mapping from CLI types → library options. +- Add deterministic **golden-output integration tests** in `crates/ewf/tests/` that: +- generate small synthetic E01/Ex01/L01/Lx01 fixtures using existing writer/test helpers +- run the Rust `ewfinfo` binary (via `std::process::Command`) and compare stdout to committed golden files for: +- default text +- `-f dfxml` +- `-F` and `-H` +- `-B` bodyfile output +- For any feature we haven’t implemented yet (e.g., extended attributes/access control entries if present in real-world files), add a test that asserts we fail with an **explicit TODO/unimplemented** marker. + +## Files most likely to change + +- [`crates/ewf/src/lib.rs`](crates/ewf/src/lib.rs) (export new ewfinfo APIs) +- [`crates/ewf/src/info.rs`](crates/ewf/src/info.rs) (keep as-is; add new full-metadata types elsewhere) +- [`crates/ewf/src/reader.rs`](crates/ewf/src/reader.rs) (expose/retain parsed metadata needed for ewfinfo) +- New: [`crates/ewf/src/ewfinfo/mod.rs`](crates/ewf/src/ewfinfo/mod.rs) +- New: [`crates/ewf/src/ewfinfo/print_text.rs`](crates/ewf/src/ewfinfo/print_text.rs) +- New: [`crates/ewf/src/ewfinfo/print_dfxml.rs`](crates/ewf/src/ewfinfo/print_dfxml.rs) +- New (preferred): `crates/ewf/src/bin/ewfinfo/` (binary crate modules, e.g. `main.rs`, `cli.rs`, `image.rs`, `logical.rs`, `bodyfile.rs`) + +## Notes / constraints + +- We’ll use libewf’s behavior/spec as reference but implement logic natively in Rust; no “silent compatibility” shims. +- Any missing surface area is left as `TODO:` + explicit `unimplemented`/`Unsupported` error (per your requirement). +- Error policy: `thiserror` in the library; `miette` at the application boundary for pretty CLI diagnostics. diff --git a/.cursor/rules/IMPORTANT-REPOWIDE.mdc b/.cursor/rules/IMPORTANT-REPOWIDE.mdc new file mode 100644 index 0000000..df6ddfb --- /dev/null +++ b/.cursor/rules/IMPORTANT-REPOWIDE.mdc @@ -0,0 +1,7 @@ +--- +alwaysApply: true +--- + +# No best-effort code. + +When commiting code to this repository - AVOID using "best-effort" code. If you do not have the means to implement some part of a solution to be SPEC EXACT - prefer leaving a TODO, or unimplemented path altogether! diff --git a/Cargo.toml b/Cargo.toml index 7a15e13..a7fe5ac 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -19,6 +19,7 @@ members = [ "crates/ewf", "crates/ntfs", "crates/ntfs-explorer-gui", + "crates/dfxml", ] [dependencies] diff --git a/crates/dfxml/Cargo.toml b/crates/dfxml/Cargo.toml new file mode 100644 index 0000000..e84128d --- /dev/null +++ b/crates/dfxml/Cargo.toml @@ -0,0 +1,11 @@ +[package] +name = "dfxml" +version = "0.1.0" +edition = "2024" + +[dependencies] +quick-xml = "0.38.4" +thiserror = "2" + +[dev-dependencies] +tempfile = "3.23" diff --git a/crates/dfxml/schema/dfxml.xsd b/crates/dfxml/schema/dfxml.xsd new file mode 100644 index 0000000..ee58dc7 --- /dev/null +++ b/crates/dfxml/schema/dfxml.xsd @@ -0,0 +1,858 @@ + + + + + + + + This is the schema file for Digital Forensics XML, version 2.0.0-beta.0. + + If you intend to use this file as a DFXML document validator, note that you will also need to download two accompanying .xsd files under the "ref" directory. The easiest way to do this is by downloading the repository as a Git clone, or by downloading the zip archive from the Github page. + + To report issues, questions, or feature requests, please either: + * File a Github issue at this repository, seeing first if it is already filed: https://github.com/dfxml-working-group/dfxml_schema + * Email the dfxml@nist.gov mailing list. If you wish to join the mailing list, send an email to dfxml-subscribe@nist.gov (no subject or message body is necessary), and a moderator will grant access. + + + + + + + (A technical aside.) The Dublin Core and XML metadata schema needs to be imported to validate with the 'xmllint' utility. To save on validation-step network transmissions, a copy is included alongside this schema, modified to also fetch the XML Schema .xsd file locally. + Ref: https://mail.gnome.org/archives/xml/2009-November/msg00022.html + + + + + + + + The schema of XML is itself imported into this document because XML schema imports are not transitive. This allows usage of special XML attributes, such as xml:lang. + + + + + + + + + + + + + + + + + + + + + + + The version of the DFXML schema to which the DFXML file adheres. + + + + + + + + + + + Allocation status of the file object. "1" implies an allocated file, according to the metadata recorded in the file's inode and reference directory entry. To support reporting deleted file content, alloc_inode and alloc_name together should be used instead of alloc, unless the subject file system has no distinction between inodes and directory entries (such as in FAT). + + + + + + Allocation status of the file object's inode metadata structure. "1" implies allocated, "0" unallocated, and recall that null is an option and will be encountered in the case of, say, a directory entry that points to a completely removed inode. + + + + + + Allocation status of the file object's name metadata structure (the directory entry). "1" implies allocated, "0" unallocated, and recall that null is an option and will be encountered in the case of, say, an inode found independently of directory walks. + + + + + + An indicator that this volume only contains fileobjects for allocated files, not recovered or deleted files. + + + + + + + + The time the file was last accesed. + + + + + + The time the file was last backed up. (Recorded in HFS+.) + + + + + + The size of the partition's block unit, in bytes. Note that a block is not necessarily a disk sector; FAT and NTFS both use clusters as their blocks. + + + + + + + + A manifest of how the environment was set when the XML generator was compiled. Note that due to restrictions of element repetitions in XML Schema 1.0's "all" and "sequence" specificiers, the build_environment definition requires children appear in the order as generated by an exemplar utility (Fiwalk for now). + + + + + + + + + + + + + + A specific location of bytes on a mass storage device. These are grouped in a byte_runs array. Child elements are one or more cryptographic hashes of the run's content. One might use this for sector-level hashes of a file's contents. + + + + + + + + + + + + + + This attribute is used to denote whether the file's contents are resident in the file metadata structure. The SleuthKit uses this to denote residency in the NTFS MFT entry, using the corresponding flags "TSK_FS_ATTR_RES" to denote a resident file, and "TSK_FS_BLOCK_FLAG_RES" for the data block. + + + + + + + + + + + + + + + + + + The command line used to invoke the program. + + + + + + The date the program was compiled. + + + + + + The compiler (if any) used to compile the program. + + + + + + + + A block of build environment and execution provenance for the XML file. Note that due to restrictions of element repetitions in XML Schema 1.0's "all" and "sequence" specifiers, the creator definition requires children appear in the order as generated by an exemplar utility (Fiwalk for now). + + + + + + + + + + + + DEPRECATED. Will be removed in DFXML v2.0.0. + + + + + + + + + The time the file was created. Sometimes called "Birth time." + + + + + + The time the file metadata were last modified. + + + + + + A disk image. + + + + + + + + + + + + + + + + + + The time the file was recorded as deleted. (Recorded in Ext2 file systems.) + + + + + + A string describing an error encountered processing an object. + + + + + + A description of the execution environment when the XML file was generated. + + + + + + + + + + + + + + + + + + + + The file name, or full known path of the file relative to the volume root. + + + + + + A file and its metadata. Byte-location information should be recorded when possible. Note that due to restrictions of element repetitions in XML Schema 1.0's "all" and "sequence" specificiers, the fileobject definition requires children appear in the order as generated by an exemplar utility (Fiwalk for now). + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + The size of the file in bytes, as reported by the file system. + + + + + + Address of first block of the file system, in bytes. This appears to be relative to the beginning of the partition; in The SleuthKit's code base, the code "->first_block" only ever appears on the left-hand side of an assignment statement when "0" is on the right-hand-side. (That is, this is always 0 in TSK-based results.) + + + + + + A numerical encoding of the file system type. The SleuthKit uses a custom enumeration of types known to the code base; the Linux kernel source code uses a different enumeration for recognized file systems. + + + + + + A human-readable string label of the file system type. Can include annotations, such as using automatic detection to determine the precise type (e.g. leaving it up to the program to distinguish FAT12 from FAT16). + + + + + + User-group identifier of the file. + + + + + + A cryptographic hash, printed in base 16 by default. + + + + + + + + DEPRECATED. No known users of this element, and little gain seen to its use. Will be removed in 2.0.0. + + + + + + + + + + The name of the host machine in which the program was executed. + + + + + + + A unique identifier for the file. It is distinct to both the input data and the process parameters of the generating tool. (It is commonly defined by incrementing a global counter in walk-encounter order of each file and directory.) This element is used instead of the XML "id" attribute and the XML Schema "ID" attribute because one might wish to preserve a fileobject's ID when it is transferred into another document (e.g. as an original_fileobject). Using either of the more elementary ID attributes would disallow preserving an ID if it was suddenly not a unique value. + + + + + + The path (absolute or relative) to the input file. Note some utilities operate on device files, some on image files, some on other DFXML files. + + + + + + The inode number (st_ino from the stat(2) system call). File systems that do not have an "Inode" may use an alternative, distinct identifier. In The SleuthKit, FAT "Inode" numbers are calculated from the directory entry's block address; NTFS's "Inode" numbers are the MFT entry address. + + + + + + The address of the last block of the file system, relative to the beginning of the partition, in bytes. As reported by file system, after to-byte conversion. Not guaranteed to be in image (for instance, in an incomplete disk image). + + + + + + The result of running libmagic to identify the file type. + + + + + + This element can be a child of the build_environment element, or the creator element. If this element is a child of build_environment, the element is about the program as it was built. If it's a child of creator, the library is about the program as it was run. + + + + + + + + + + + The file to which a soft link refers. + + + + + + A numeric encoding of the general file type - regular, directory, soft link, etc. Numeric values are particular to The SleuthKit; the name_type element renders the values to short string representations. + + + + + + Document metadata not already defined by DFXML structures. Originally, Dublin Core was the only expected namespace. Other metadata not best expressed with Dublin Core can be brought in under their own namespaces, but only Dublin Core will be checked for valid structure. + + + + + + + + + + + File opening mode. This is the inode mode in POSIX file systems, and an encoding of various NTFS file attributes when created by The SleuthKit libraries. Recorded in base 10. + + + + + + The time the file data were last modified. + + + + + + A string representation of the general file type - regular, directory, soft link, etc. + + + + + + Unknown type + + + + + Named pipe + + + + + Character device + + + + + Directory + + + + + Block device + + + + + Regular file + + + + + Symbolic (soft) link + + + + + Socket + + + + + Shadow inode (Solaris) + + + + + Whiteout (OpenBSD) + + + + + Special (Used in The SleuthKit for added "Virtual" directories, e.g. $OrphanFiles) + + + + + Special (Used in The SleuthKit for added "Virtual" files, e.g. $FAT1) + + + + + + + + + The number of hard links to this file's inode. + + + + + + A file lacking a referencing metadata structure. + + + + + + The operating system release (reported by uname -r). + + + + + + The operating system name (reported by uname -s). + + + + + + The operating system version (reported by uname -v). + + + + + + A container for an element referencing the parent of this fileobject. The inode child specified here would contain the inode number of the parent directory. The root directory of a file system should not have a parent fileobject. + + + + + + + + + + + + The partition in which the file resides. 1-based counter of the partition order. + + + + + + The index value for this partition within the containing partition table. Many partition systems use a numeric value, but some (such as the BSD Disk Label) use a character value. + + + + + + The offset of the partition from the beginning of the image file, in bytes. + + + + + + A disk partition. To represent MBR extended partition contents, a partition object can nest within another partition object. + + + + + + + + + + + + + + + + + + + + A partition system. + + + + + + + + + + + + + + + + + The name of the XML-generating program. + + + + + + A human-readable string label of the partition system type. + + + + + + The numerical encoding of the partition's type, as recorded in the containing partition table. + + + + + + A human-readable string label of the partition type. + + + + + + + This element encodes a "rusage" C structure, as provided by the "getrusage" function after the file walk is complete. In addition to the "rusage" fields, the element may include an element for elapsed wall clock time in seconds. + + + + + + + + + + + + + + + + + + + + The size of a disk sector in this volume. Note that this is not necessarily the same unit as the volume will use for its blocks (see block_size element). + + + + + + The NTFS sequence number. + + + + + + + + + + + + + + + The date and time that the program was executed. + + + + + + The user id that owns a file. Numeric in POSIX file systems, but a string for SIDs in NTFS. For SID reference, see: http://www.ntfs.com/ntfs-permissions-security-identifier.htm. + + + + + + DEPRECATED. See alloc. + + + + + + This file's metadata structure has never been used (had an attribute populated), or possibly never been allocated. + + + + + + + This file's metadata structure has at least one attribute populated. + + + + + + The username under which the program was executed. + + + + + + The version of the XML-generating program. + + + + + + A mass storage system volume, which is defined as a collection of byte blocks that are all the same size. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + A 0-or-1 Boolean value. + + + + + + + + + + A general structure to represent a xs:dateTime with the precision attribute. + + + + + + The precision of this timestamp, in seconds. + + + + + + + + + + The hash algorithm that applies to this object. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/crates/dfxml/schema/ref/dc.xsd b/crates/dfxml/schema/ref/dc.xsd new file mode 100644 index 0000000..82239ee --- /dev/null +++ b/crates/dfxml/schema/ref/dc.xsd @@ -0,0 +1,119 @@ + + + + + + DCMES 1.1 XML Schema + XML Schema for http://purl.org/dc/elements/1.1/ namespace + + Created 2008-02-11 + + Created by + + Tim Cole (t-cole3@uiuc.edu) + Tom Habing (thabing@uiuc.edu) + Jane Hunter (jane@dstc.edu.au) + Pete Johnston (p.johnston@ukoln.ac.uk), + Carl Lagoze (lagoze@cs.cornell.edu) + + This schema declares XML elements for the 15 DC elements from the + http://purl.org/dc/elements/1.1/ namespace. + + It defines a complexType SimpleLiteral which permits mixed content + and makes the xml:lang attribute available. It disallows child elements by + use of minOcccurs/maxOccurs. + + However, this complexType does permit the derivation of other complexTypes + which would permit child elements. + + All elements are declared as substitutable for the abstract element any, + which means that the default type for all elements is dc:SimpleLiteral. + + + + + + + + + + + + + + This is the default type for all of the DC elements. + It permits text content only with optional + xml:lang attribute. + Text is allowed because mixed="true", but sub-elements + are disallowed because minOccurs="0" and maxOccurs="0" + are on the xs:any tag. + + This complexType allows for restriction or extension permitting + child elements. + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + This group is included as a convenience for schema authors + who need to refer to all the elements in the + http://purl.org/dc/elements/1.1/ namespace. + + + + + + + + + + + + + + This complexType is included as a convenience for schema authors who need to define a root + or container element for all of the DC elements. + + + + + + + + + + diff --git a/crates/dfxml/schema/ref/xml.xsd b/crates/dfxml/schema/ref/xml.xsd new file mode 100644 index 0000000..8a803be --- /dev/null +++ b/crates/dfxml/schema/ref/xml.xsd @@ -0,0 +1,146 @@ + + + + + + See http://www.w3.org/XML/1998/namespace.html and + http://www.w3.org/TR/REC-xml for information about this namespace. + + This schema document describes the XML namespace, in a form + suitable for import by other schema documents. + + Note that local names in this namespace are intended to be defined + only by the World Wide Web Consortium or its subgroups. The + following names are currently defined in this namespace and should + not be used with conflicting semantics by any Working Group, + specification, or document instance: + + base (as an attribute name): denotes an attribute whose value + provides a URI to be used as the base for interpreting any + relative URIs in the scope of the element on which it + appears; its value is inherited. This name is reserved + by virtue of its definition in the XML Base specification. + + id (as an attribute name): denotes an attribute whose value + should be interpreted as if declared to be of type ID. + The xml:id specification is not yet a W3C Recommendation, + but this attribute is included here to facilitate experimentation + with the mechanisms it proposes. Note that it is _not_ included + in the specialAttrs attribute group. + + lang (as an attribute name): denotes an attribute whose value + is a language code for the natural language of the content of + any element; its value is inherited. This name is reserved + by virtue of its definition in the XML specification. + + space (as an attribute name): denotes an attribute whose + value is a keyword indicating what whitespace processing + discipline is intended for the content of the element; its + value is inherited. This name is reserved by virtue of its + definition in the XML specification. + + Father (in any context at all): denotes Jon Bosak, the chair of + the original XML Working Group. This name is reserved by + the following decision of the W3C XML Plenary and + XML Coordination groups: + + In appreciation for his vision, leadership and dedication + the W3C XML Plenary on this 10th day of February, 2000 + reserves for Jon Bosak in perpetuity the XML name + xml:Father + + + + + This schema defines attributes and an attribute group + suitable for use by + schemas wishing to allow xml:base, xml:lang, xml:space or xml:id + attributes on elements they define. + + To enable this, such a schema must import this schema + for the XML namespace, e.g. as follows: + <schema . . .> + . . . + <import namespace="http://www.w3.org/XML/1998/namespace" + schemaLocation="http://www.w3.org/2005/08/xml.xsd"/> + + Subsequently, qualified reference to any of the attributes + or the group defined below will have the desired effect, e.g. + + <type . . .> + . . . + <attributeGroup ref="xml:specialAttrs"/> + + will define a type which will schema-validate an instance + element with any of those attributes + + + + In keeping with the XML Schema WG's standard versioning + policy, this schema document will persist at + http://www.w3.org/2005/08/xml.xsd. + At the date of issue it can also be found at + http://www.w3.org/2001/xml.xsd. + The schema document at that URI may however change in the future, + in order to remain compatible with the latest version of XML Schema + itself, or with the XML namespace itself. In other words, if the XML + Schema or XML namespaces change, the version of this document at + http://www.w3.org/2001/xml.xsd will change + accordingly; the version at + http://www.w3.org/2005/08/xml.xsd will not change. + + + + + + Attempting to install the relevant ISO 2- and 3-letter + codes as the enumerated possible values is probably never + going to be a realistic possibility. See + RFC 3066 at http://www.ietf.org/rfc/rfc3066.txt and the IANA registry + at http://www.iana.org/assignments/lang-tag-apps.htm for + further information. + + The union allows for the 'un-declaration' of xml:lang with + the empty string. + + + + + + + + + + + + + + + + + + + + + + + + See http://www.w3.org/TR/xmlbase/ for + information about this attribute. + + + + + + See http://www.w3.org/TR/xml-id/ for + information about this attribute. + + + + + + + + + + diff --git a/crates/dfxml/src/lib.rs b/crates/dfxml/src/lib.rs new file mode 100644 index 0000000..1967727 --- /dev/null +++ b/crates/dfxml/src/lib.rs @@ -0,0 +1,515 @@ +//! Digital Forensics XML (DFXML) types and writers. +//! +//! This crate provides a small, schema-aligned subset of DFXML for use by the workspace tools. +//! Output is written using `quick-xml` to ensure well-formed XML and correct escaping. +//! +//! ## Schema reference +//! +//! This crate targets DFXML schema **2.0.0-beta.0** as published by the DFXML Working Group: +//! +//! - Reference repo (pinned): `external/refs/repos/dfxml-working-group__dfxml_schema.commit` +//! - Upstream: `https://github.com/dfxml-working-group/dfxml_schema` +//! - Local clone (for convenience): `external/refs/repos/dfxml-working-group__dfxml_schema/` +//! +//! The schema defines the namespace: +//! +//! - `http://www.forensicswiki.org/wiki/Category:Digital_Forensics_XML` +//! +//! and requires a `` root element with a `version` attribute, and a `` child. +//! +//! ## Scope (intentionally small) +//! +//! We only implement the subset needed by this repository’s tools today: +//! +//! - `` root +//! - `` with Dublin Core elements (e.g. `dc:type`) +//! - `` (program/version/build_environment/execution_environment) +//! - `` (image_filename) +//! - `` with `` and `/` +//! - `...` +//! +//! Missing parts of the DFXML schema are not “best-effort”. Instead, callers should treat absent +//! types as “not supported yet” and either omit them or model them explicitly in this crate. + +use std::borrow::Cow; + +use quick_xml::Writer; +use quick_xml::events::{BytesDecl, BytesEnd, BytesStart, BytesText, Event}; + +/// DFXML XML namespace URI. +pub const DFXML_NS: &str = "http://www.forensicswiki.org/wiki/Category:Digital_Forensics_XML"; + +/// DFXML schema version targeted by this crate. +pub const DFXML_SCHEMA_VERSION: &str = "2.0.0-beta.0"; + +/// Dublin Core namespace URI. +pub const DC_NS: &str = "http://purl.org/dc/elements/1.1/"; + +#[derive(Debug, thiserror::Error)] +pub enum DfxmlError { + #[error(transparent)] + Io(#[from] std::io::Error), + #[error(transparent)] + Utf8(#[from] std::string::FromUtf8Error), +} + +pub type Result = std::result::Result; + +/// A DFXML document. +#[derive(Debug, Clone, Default)] +pub struct DfxmlDocument { + /// The schema version string placed in the root `dfxml@version` attribute. + pub schema_version: Cow<'static, str>, + + pub metadata: Metadata, + pub creator: Option, + pub source: Option, + pub diskimageobjects: Vec, +} + +impl DfxmlDocument { + pub fn new() -> Self { + Self { + schema_version: Cow::Borrowed(DFXML_SCHEMA_VERSION), + metadata: Metadata::default(), + creator: None, + source: None, + diskimageobjects: Vec::new(), + } + } + + pub fn to_xml_string(&self) -> Result { + let mut w = Writer::new_with_indent(Vec::new(), b'\t', 1); + + w.write_event(Event::Decl(BytesDecl::new("1.0", Some("UTF-8"), None)))?; + w.write_event(Event::Text(BytesText::new("\n")))?; + + let mut root = BytesStart::new("dfxml"); + root.push_attribute(("xmlns", DFXML_NS)); + root.push_attribute(("xmlns:dc", DC_NS)); + root.push_attribute(("version", self.schema_version.as_ref())); + w.write_event(Event::Start(root))?; + + self.metadata.write(&mut w)?; + + if let Some(c) = &self.creator { + c.write(&mut w)?; + } + if let Some(s) = &self.source { + s.write(&mut w)?; + } + for d in &self.diskimageobjects { + d.write(&mut w)?; + } + + w.write_event(Event::End(BytesEnd::new("dfxml")))?; + w.write_event(Event::Text(BytesText::new("\n")))?; + + Ok(String::from_utf8(w.into_inner())?) + } +} + +/// DFXML metadata container. +/// +/// In the schema, `` is an “any”-container; this crate focuses on well-known Dublin Core +/// elements. +#[derive(Debug, Clone, Default)] +pub struct Metadata { + pub dublin_core: Vec, +} + +impl Metadata { + fn write(&self, w: &mut Writer>) -> Result<()> { + if self.dublin_core.is_empty() { + w.write_event(Event::Empty(BytesStart::new("metadata")))?; + return Ok(()); + } + + w.write_event(Event::Start(BytesStart::new("metadata")))?; + for e in &self.dublin_core { + e.write(w)?; + } + w.write_event(Event::End(BytesEnd::new("metadata")))?; + Ok(()) + } +} + +#[derive(Debug, Clone)] +pub struct DublinCoreElement { + pub name: DublinCoreElementName, + pub value: String, +} + +impl DublinCoreElement { + fn write(&self, w: &mut Writer>) -> Result<()> { + // Write as a namespaced element (e.g. ``). + let tag = format!("dc:{}", self.name.as_str()); + w.write_event(Event::Start(BytesStart::new(tag.as_str())))?; + w.write_event(Event::Text(BytesText::new(&self.value)))?; + w.write_event(Event::End(BytesEnd::new(tag.as_str())))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum DublinCoreElementName { + Title, + Creator, + Subject, + Description, + Publisher, + Contributor, + Date, + Type, + Format, + Identifier, + Source, + Language, + Relation, + Coverage, + Rights, +} + +impl DublinCoreElementName { + pub fn as_str(self) -> &'static str { + match self { + DublinCoreElementName::Title => "title", + DublinCoreElementName::Creator => "creator", + DublinCoreElementName::Subject => "subject", + DublinCoreElementName::Description => "description", + DublinCoreElementName::Publisher => "publisher", + DublinCoreElementName::Contributor => "contributor", + DublinCoreElementName::Date => "date", + DublinCoreElementName::Type => "type", + DublinCoreElementName::Format => "format", + DublinCoreElementName::Identifier => "identifier", + DublinCoreElementName::Source => "source", + DublinCoreElementName::Language => "language", + DublinCoreElementName::Relation => "relation", + DublinCoreElementName::Coverage => "coverage", + DublinCoreElementName::Rights => "rights", + } + } +} + +/// DFXML creator block. +#[derive(Debug, Clone, Default)] +pub struct Creator { + pub program: Option, + pub version: Option, + pub build_environment: Option, + pub execution_environment: Option, +} + +impl Creator { + fn write(&self, w: &mut Writer>) -> Result<()> { + w.write_event(Event::Start(BytesStart::new("creator")))?; + + if let Some(s) = &self.program { + write_text_element(w, "program", s)?; + } + if let Some(s) = &self.version { + write_text_element(w, "version", s)?; + } + if let Some(be) = &self.build_environment { + be.write(w)?; + } + if let Some(ee) = &self.execution_environment { + ee.write(w)?; + } + + w.write_event(Event::End(BytesEnd::new("creator")))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Default)] +pub struct BuildEnvironment { + pub compiler: Option, + /// ISO-8601 xs:dateTime. If unknown, leave `None` (schema type is dateTime). + pub compilation_date: Option, + pub libraries: Vec, +} + +impl BuildEnvironment { + fn write(&self, w: &mut Writer>) -> Result<()> { + w.write_event(Event::Start(BytesStart::new("build_environment")))?; + + if let Some(s) = &self.compiler { + write_text_element(w, "compiler", s)?; + } + if let Some(s) = &self.compilation_date { + write_text_element(w, "compilation_date", s)?; + } + for lib in &self.libraries { + lib.write(w)?; + } + + w.write_event(Event::End(BytesEnd::new("build_environment")))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Default)] +pub struct ExecutionEnvironment { + pub os_sysname: Option, + pub arch: Option, +} + +impl ExecutionEnvironment { + fn write(&self, w: &mut Writer>) -> Result<()> { + w.write_event(Event::Start(BytesStart::new("execution_environment")))?; + if let Some(s) = &self.os_sysname { + write_text_element(w, "os_sysname", s)?; + } + if let Some(s) = &self.arch { + write_text_element(w, "arch", s)?; + } + w.write_event(Event::End(BytesEnd::new("execution_environment")))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Default)] +pub struct Library { + pub name: Option, + pub version: Option, +} + +impl Library { + fn write(&self, w: &mut Writer>) -> Result<()> { + let mut el = BytesStart::new("library"); + if let Some(name) = &self.name { + el.push_attribute(("name", name.as_str())); + } + if let Some(ver) = &self.version { + el.push_attribute(("version", ver.as_str())); + } + w.write_event(Event::Empty(el))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Default)] +pub struct Source { + pub image_filenames: Vec, +} + +impl Source { + fn write(&self, w: &mut Writer>) -> Result<()> { + if self.image_filenames.is_empty() { + return Ok(()); + } + w.write_event(Event::Start(BytesStart::new("source")))?; + for p in &self.image_filenames { + write_text_element(w, "image_filename", p)?; + } + w.write_event(Event::End(BytesEnd::new("source")))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Default)] +pub struct DiskImageObject { + pub sector_size: Option, + pub byte_runs: Vec, +} + +impl DiskImageObject { + fn write(&self, w: &mut Writer>) -> Result<()> { + w.write_event(Event::Start(BytesStart::new("diskimageobject")))?; + + if !self.byte_runs.is_empty() { + w.write_event(Event::Start(BytesStart::new("byte_runs")))?; + for br in &self.byte_runs { + br.write(w)?; + } + w.write_event(Event::End(BytesEnd::new("byte_runs")))?; + } + + if let Some(sector_size) = self.sector_size { + write_text_element(w, "sector_size", §or_size.to_string())?; + } + + w.write_event(Event::End(BytesEnd::new("diskimageobject")))?; + Ok(()) + } +} + +#[derive(Debug, Clone)] +pub struct ByteRun { + pub img_offset: u64, + pub len: u64, + /// Optional `byte_run@type` attribute (string in schema). + pub kind: Option, + pub hashdigests: Vec, +} + +impl ByteRun { + fn write(&self, w: &mut Writer>) -> Result<()> { + let mut el = BytesStart::new("byte_run"); + let off_s = self.img_offset.to_string(); + let len_s = self.len.to_string(); + el.push_attribute(("img_offset", off_s.as_str())); + el.push_attribute(("len", len_s.as_str())); + if let Some(kind) = &self.kind { + el.push_attribute(("type", kind.as_str())); + } + + if self.hashdigests.is_empty() { + w.write_event(Event::Empty(el))?; + return Ok(()); + } + + w.write_event(Event::Start(el))?; + for h in &self.hashdigests { + h.write(w)?; + } + w.write_event(Event::End(BytesEnd::new("byte_run")))?; + Ok(()) + } +} + +#[derive(Debug, Clone)] +pub struct HashDigest { + pub algorithm: HashDigestType, + pub value: String, +} + +impl HashDigest { + fn write(&self, w: &mut Writer>) -> Result<()> { + let mut el = BytesStart::new("hashdigest"); + el.push_attribute(("type", self.algorithm.as_str())); + w.write_event(Event::Start(el))?; + w.write_event(Event::Text(BytesText::new(&self.value)))?; + w.write_event(Event::End(BytesEnd::new("hashdigest")))?; + Ok(()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum HashDigestType { + Md5, + Sha1, +} + +impl HashDigestType { + pub fn as_str(self) -> &'static str { + match self { + HashDigestType::Md5 => "md5", + HashDigestType::Sha1 => "sha1", + } + } +} + +fn write_text_element(w: &mut Writer>, name: &str, text: &str) -> std::io::Result<()> { + w.write_event(Event::Start(BytesStart::new(name)))?; + w.write_event(Event::Text(BytesText::new(text)))?; + w.write_event(Event::End(BytesEnd::new(name)))?; + Ok(()) +} + +#[cfg(test)] +mod tests { + use super::*; + use std::path::Path; + use std::process::Command; + + #[test] + fn test_minimal_dfxml_shape_matches_schema_example() { + let doc = DfxmlDocument::new(); + let xml = doc.to_xml_string().unwrap(); + + // Must have correct namespace and required version attribute. + assert!(xml.contains("") || xml.contains("")); + } + + #[test] + fn test_diskimageobject_byte_run_and_hashdigest_shape() { + let mut doc = DfxmlDocument::new(); + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Type, + value: "Disk Image".to_string(), + }); + doc.creator = Some(Creator { + program: Some("ewfinfo".to_string()), + version: Some("0.1.0".to_string()), + build_environment: None, + execution_environment: Some(ExecutionEnvironment { + os_sysname: Some("macos".to_string()), + arch: Some("aarch64".to_string()), + }), + }); + doc.source = Some(Source { + image_filenames: vec!["image.E01".to_string()], + }); + doc.diskimageobjects.push(DiskImageObject { + sector_size: Some(512), + byte_runs: vec![ByteRun { + img_offset: 0, + len: 1474560, + kind: Some("image".to_string()), + hashdigests: vec![ + HashDigest { + algorithm: HashDigestType::Md5, + value: "deadbeef".to_string(), + }, + HashDigest { + algorithm: HashDigestType::Sha1, + value: "cafebabe".to_string(), + }, + ], + }], + }); + + let xml = doc.to_xml_string().unwrap(); + assert!(xml.contains("")); + assert!(xml.contains("512")); + assert!(xml.contains("")); + assert!(xml.contains("img_offset=\"0\"")); + assert!(xml.contains("len=\"1474560\"")); + assert!(xml.contains("type=\"image\"")); + assert!(xml.contains("deadbeef")); + assert!(xml.contains("cafebabe")); + } + + #[test] + fn test_xmllint_validates_output_against_vendored_schema_if_available() { + // `xmllint` is available on macOS/Linux, but not guaranteed on Windows CI. + let Ok(v) = Command::new("xmllint").arg("--version").output() else { + return; + }; + if !v.status.success() { + return; + } + + let doc = DfxmlDocument::new(); + let xml = doc.to_xml_string().unwrap(); + + let dir = tempfile::tempdir().unwrap(); + let xml_path = dir.path().join("out.dfxml"); + std::fs::write(&xml_path, xml).unwrap(); + + let schema_path = Path::new(env!("CARGO_MANIFEST_DIR")).join("schema/dfxml.xsd"); + + let out = Command::new("xmllint") + .arg("--noout") + .arg("--schema") + .arg(schema_path) + .arg(&xml_path) + .output() + .unwrap(); + + assert!( + out.status.success(), + "xmllint failed: {}\n{}", + String::from_utf8_lossy(&out.stdout), + String::from_utf8_lossy(&out.stderr) + ); + } +} diff --git a/crates/ewf/.cursor/rules/IMPORTANT.mdc b/crates/ewf/.cursor/rules/IMPORTANT.mdc new file mode 100644 index 0000000..182b614 --- /dev/null +++ b/crates/ewf/.cursor/rules/IMPORTANT.mdc @@ -0,0 +1,81 @@ +--- +alwaysApply: true +--- + +# Spec driven development + +When writing code or debugging issues, YOU MUST compare with original libewf implementation AND SPEC. +AVOID using web search/firecrawl to find sources from libewf! you must use local sources. + +If you are running in a worktree - you can clone the libewf as specified by the commit hash in the external/refs/repos/libyal__libewf.commit file. + +When accessing resources from libewf local sources, your built- grep tool may not be able to find the sources. +Instead prefer using ripgrep (--no-ignore) to search the local sources. + +Reference: +- libewf implementation: `external/libewf/` +- libewf specification: + - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` + - `external/libewf/documentation/Expert Witness Compression Format 2 (EWF2).asciidoc` + +# Rust style + +Prefer storing structs in: + + - 'src/ewf1/...' + - 'src/ewf2/...' + +generally: + - modules should have documentation that explains reference to spec. + - all public functions, and struct members should have documentation outline that explain their purpose, and provenance. + - prefer attaching methods/parsing to their structs, rather than having "c-style" free floating functions. + + +## Binary parsing + +We strive to maintain symmetry between the reader and writer implementations. +We use binrw to parse and write binary data. + +DO use binrw methods to make parser declerative and self describing. +``` +/// EWF2 container kind (Ex01 vs Lx01). +/// +/// The kind is encoded in the file header signature. +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf2Kind { + /// EWF2-Ex01 (`EVF2\r\n\x81\x00`) + #[brw(magic = b"EVF2\r\n\x81\0")] + Ex01, + /// EWF2-Lx01 (`LEF2\r\n\x81\x00`) + #[brw(magic = b"LEF2\r\n\x81\0")] + Lx01, +} + +impl Ewf2Kind { + pub(crate) fn signature(self) -> [u8; 8] { + match self { + Self::Ex01 => EWF2_EVF_SIGNATURE, + Self::Lx01 => EWF2_LEF_SIGNATURE, + } + } +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) struct Ewf2FileHeader { + pub(crate) kind: Ewf2Kind, + #[br(assert( + major == EWF2_VERSION_MAJOR, + "unsupported EWF2 major version: {}", + major + ))] + pub(crate) major: u8, + pub(crate) minor: u8, + pub(crate) compression_method: u16, + pub(crate) segment_number: u32, + pub(crate) set_id: [u8; 16], +} +``` diff --git a/crates/ewf/COMPAT.md b/crates/ewf/COMPAT.md index 59a2e4f..20f8087 100644 --- a/crates/ewf/COMPAT.md +++ b/crates/ewf/COMPAT.md @@ -19,7 +19,7 @@ Reference commit pinned in this repository: - [x] Uncompressed chunks - [x] Chunk Adler32 verification - [ ] EnCase 5/6 metadata edge-cases (header/header2 variants) — partial -- [ ] EWF1 “sessions” section (optical media) — not implemented +- [x] EWF1 “sessions” section (optical media) — run parsing (used for `ewfinfo` output) ### EWF2 (EVF2) — `.Ex01` @@ -80,3 +80,35 @@ Reference commit pinned in this repository: - [x] AccessData “ADCRYPT” container (FTK/AD encryption) — explicitly rejected with a clear error +## `ewfinfo` CLI parity (this repository) + +This section tracks parity with libewf’s `ewfinfo` tool behavior (`external/libewf/ewftools/ewfinfo.c`, +`external/libewf/ewftools/info_handle.c`, `external/libewf/manuals/ewfinfo.1`). + +Notes: +- **Library support**: spec-oriented metadata extraction lives in `crates/ewf/src/metadata.rs` and + `EwfReader::image_metadata()`. +- **Image-mode reporting/rendering** is binary-owned: `crates/ewf/src/bin/ewfinfo/ewfinfo/`. +- **Logical evidence outputs** (`-F`/`-H`/`-B`) are implemented in the `ewfinfo` binary target only: + `crates/ewf/src/bin/ewfinfo/`. + +### Image mode (`.E01` / `.S01` / `.Ex01`) — metadata report + +- [x] Text output (default / `-f text`) +- [x] DFXML output (`-f dfxml`) — **schema-aligned ``** (DFXML 2.0.0-beta.0) via `crates/dfxml` +- [x] Section filtering: `-i` (acquiry only), `-m` (media+hashes+sessions+tracks), `-e` (errors only) +- [x] Date formatting `-d ctime|dm|md|iso8601` for acquisition/system date header values +- [x] `-A ascii` (EWF1 header decoding) +- [ ] `-A windows-*` codepages for EWF1 header decoding (explicit `Unsupported` for now) +- [ ] `-v` verbose output parity (flag is accepted, but we don’t emit libewf-style verbose traces yet) +- [ ] libewf “DFXML” (`ewfobjects` root + `ewfinfo` sections) output compatibility — we intentionally emit schema-aligned DFXML instead + +### Logical evidence mode (`.L01` / `.Lx01`) — tree + bodyfile + +- [x] `-H` logical files hierarchy (text) +- [x] `-F ` file entry info (text) — **subset** of libewf fields +- [x] `-B ` bodyfile output (Sleuthkit 3.x+ columns) — **subset** of libewf behavior +- [x] `-s /|\\` path separator for text/hierarchy/bodyfile name fields +- [ ] `-f dfxml` for logical modes (`-H`/`-F`/`-B`) — not implemented (test asserts this) +- [ ] Full file-entry metadata parity (ACLs, owners/groups, short name, etc.) + diff --git a/crates/ewf/Cargo.toml b/crates/ewf/Cargo.toml index c57f6e6..9f0b0de 100644 --- a/crates/ewf/Cargo.toml +++ b/crates/ewf/Cargo.toml @@ -20,6 +20,14 @@ lru = "0.16.1" md-5 = "0.10" sha1 = "0.10" rand = "0.9" +clap = { version = "4", features = ["derive"] } +miette = { version = "7", features = ["fancy"] } +dfxml = { path = "../dfxml" } +jiff = { version = "0.2.16", default-features = false, features = ["std", "alloc", "tz-system", "tzdb-zoneinfo"] } +textwrap = "0.16.2" +quick-xml = "0.38.4" +binrw = "0.15.0" +bitflags = "2.10.0" [dev-dependencies] doc-comment = "0.3" diff --git a/crates/ewf/src/bin/ewfinfo/bodyfile.rs b/crates/ewf/src/bin/ewfinfo/bodyfile.rs new file mode 100644 index 0000000..1887bcc --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/bodyfile.rs @@ -0,0 +1,119 @@ +//! Sleuthkit bodyfile output for `ewfinfo -B`. +//! +//! libewf’s `ewfinfo` can emit logical file information in “bodyfile” format (`-B`). The bodyfile +//! is written to a file path provided on the command line, and (in libewf) is driven by traversing +//! the logical file tree (`-H`) or printing a specific entry (`-F`). +//! +//! References: +//! - `external/libewf/ewftools/info_handle.c` (bodyfile columns and traversal) +//! - `external/libewf/ewftools/bodyfile.c` (escaping rules for the name field) +//! - `external/libewf/manuals/ewfinfo.1` + +use ewf::LefEntry; + +use crate::logical; + +/// Escapes a bodyfile name value. +/// +/// This is a small, Rust-native equivalent of libewf’s +/// `bodyfile_path_string_copy_from_file_entry_path(...)`: +/// +/// - Escape character (`\`) becomes `\\` +/// - Bodyfile separator (`|`) becomes `\|` +/// - ASCII control characters are replaced by `\x##` (lowercase hex) +pub fn escape_bodyfile_name(value: &str) -> String { + let mut out = String::new(); + for ch in value.chars() { + let code = ch as u32; + + // Control characters. + if code <= 0x1f || (0x7f..=0x9f).contains(&code) { + out.push('\\'); + out.push('x'); + out.push_str(&format!("{code:02x}")); + continue; + } + + // Escape `\` and `|`. + if ch == '\\' || ch == '|' { + out.push('\\'); + out.push(ch); + continue; + } + + out.push(ch); + } + out +} + +pub fn render_bodyfile_line(entry: &LefEntry, separator: char) -> String { + // Colums in a Sleuthkit 3.x and later bodyfile (as used by libewf): + // MD5|name|inode|mode_as_string|UID|GID|size|atime|mtime|ctime|crtime + // + // libewf currently prints: + // - a constant `0` as MD5, + // - `0` for UID/GID (TODO in upstream), + // - seconds since epoch as `%.9f`. + let name = escape_bodyfile_name(&logical::display_path(&entry.path, separator)); + let inode = entry.file_identifier.unwrap_or(0); + + let mode = if entry.is_dir { + "drwxrwxrwx" + } else { + "-rwxrwxrwx" + }; + + let atime = entry.access_time.unwrap_or(0) as f64; + let mtime = entry.modification_time.unwrap_or(0) as f64; + let ctime = entry.entry_modification_time.unwrap_or(0) as f64; + let crtime = entry.creation_time.unwrap_or(0) as f64; + + format!( + "0|{name}|{inode}|{mode}|0|0|{}|{atime:.9}|{mtime:.9}|{ctime:.9}|{crtime:.9}\n", + entry.size + ) +} + +pub fn render_bodyfile(entries: &[LefEntry], separator: char) -> String { + let mut out = String::new(); + for e in entries { + out.push_str(&render_bodyfile_line(e, separator)); + } + out +} + +#[cfg(test)] +mod tests { + use super::*; + use ewf::LefEntry; + + #[test] + fn test_escape_bodyfile_name_escapes_separator_and_backslash() { + assert_eq!(escape_bodyfile_name(r"dir\file|name"), r"dir\\file\|name"); + } + + #[test] + fn test_escape_bodyfile_name_escapes_control_chars() { + assert_eq!(escape_bodyfile_name("a\nb"), r"a\x0ab"); + } + + #[test] + fn test_render_bodyfile_line_shape() { + let e = LefEntry { + path: "dir/file.txt".to_string(), + is_dir: false, + size: 5, + extents: vec![], + file_identifier: Some(42), + access_time: Some(100), + modification_time: Some(200), + entry_modification_time: Some(300), + creation_time: Some(400), + }; + + assert_eq!( + render_bodyfile_line(&e, '/'), + "0|dir/file.txt|42|-rwxrwxrwx|0|0|5|100.000000000|200.000000000|300.000000000|400.000000000\n" + ); + } +} diff --git a/crates/ewf/src/bin/ewfinfo/cli.rs b/crates/ewf/src/bin/ewfinfo/cli.rs new file mode 100644 index 0000000..beb6003 --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/cli.rs @@ -0,0 +1,385 @@ +//! clap-driven CLI parsing for the `ewfinfo` binary. +//! +//! This module is deliberately **binary-only**: it translates libewf-compatible command-line flags +//! into calls to the `ewf` crate. The `ewf` library provides spec-oriented metadata extraction, +//! while this binary owns the reporting/rendering layer and user-facing diagnostics. +//! +//! References: +//! - `external/libewf/ewftools/ewfinfo.c` (option parsing + modes) +//! - `external/libewf/manuals/ewfinfo.1` (flag documentation) + +use std::io::Write as _; +use std::path::PathBuf; + +use clap::{ArgAction, Parser, ValueEnum}; +use miette::{Context as _, IntoDiagnostic as _, miette}; + +use ewf::EwfFormat; +use ewf::EwfReader; +use ewf::LefReader; + +use crate::ewfinfo::{ + EwfInfoColorMode, EwfInfoDateFormat, EwfInfoPrintOptions, EwfInfoReport, EwfInfoSections, + HeaderCodepage, +}; +use crate::{bodyfile, logical}; + +/// Show meta data stored in EWF files. +#[derive(Debug, Parser)] +#[command(name = "ewfinfo", disable_version_flag = true)] +pub struct Cli { + /// The codepage of the header section, options: ascii (default), windows-874, windows-932, + /// windows-936, windows-949, windows-950, windows-1250, windows-1251, windows-1252, + /// windows-1253, windows-1254, windows-1255, windows-1256, windows-1257 or windows-1258 + #[arg(short = 'A', value_enum, value_name = "codepage", default_value_t = HeaderCodepageArg::Ascii)] + pub header_codepage: HeaderCodepageArg, + + /// Output logical files information as a bodyfile. + #[arg(short = 'B', value_name = "bodyfile")] + pub bodyfile: Option, + + /// Show information about a specific file entry path. + #[arg(short = 'F', value_name = "file_entry", conflicts_with = "hierarchy")] + pub file_entry: Option, + + /// The date format, options: ctime (default), dm (day/month), md (month/day), iso8601. + #[arg(short = 'd', value_enum, value_name = "date_format", default_value_t = DateFormatArg::Ctime)] + pub date_format: DateFormatArg, + + /// Only show EWF read error information. + #[arg(short = 'e', action = ArgAction::SetTrue, conflicts_with_all = ["acquiry_only", "media_only"])] + pub errors_only: bool, + + /// Shows the logical files hierarchy. + #[arg(short = 'H', action = ArgAction::SetTrue)] + pub hierarchy: bool, + + /// Only show EWF acquiry information. + #[arg(short = 'i', action = ArgAction::SetTrue, conflicts_with_all = ["errors_only", "media_only"])] + pub acquiry_only: bool, + + /// Only show EWF media information. + #[arg(short = 'm', action = ArgAction::SetTrue, conflicts_with_all = ["errors_only", "acquiry_only"])] + pub media_only: bool, + + /// Path segment separator, options: `/` (default), `\\`. + #[arg(short = 's', value_enum, value_name = "separator", default_value_t = PathSeparator::Slash)] + pub separator: PathSeparator, + + /// Specify the output format, options: text (default), dfxml. + #[arg(short = 'f', value_enum, value_name = "format", default_value_t = OutputFormat::Text)] + pub format: OutputFormat, + + /// Verbose output to stderr. + #[arg(short = 'v', action = ArgAction::SetTrue)] + pub verbose: bool, + + /// Print version and exit. + #[arg(short = 'V', long = "version", action = ArgAction::SetTrue)] + pub version: bool, + + /// The first or the entire set of EWF segment files. + #[arg(value_name = "ewf_files", required = true)] + pub ewf_files: Vec, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, ValueEnum)] +pub enum OutputFormat { + /// Human-readable text output. + #[value(name = "text")] + Text, + /// DFXML output (schema-aligned DFXML 2.0.0-beta.0; not yet implemented for logical-evidence modes). + #[value(name = "dfxml")] + Dfxml, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, ValueEnum)] +pub enum PathSeparator { + /// `/` (default). + #[value(name = "/")] + Slash, + /// `\\` (Windows-style). + #[value(name = "\\")] + Backslash, +} + +impl PathSeparator { + pub fn as_char(self) -> char { + match self { + Self::Slash => '/', + Self::Backslash => '\\', + } + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, ValueEnum)] +pub enum HeaderCodepageArg { + #[value(name = "ascii")] + Ascii, + #[value(name = "windows-874")] + Windows874, + #[value(name = "windows-932")] + Windows932, + #[value(name = "windows-936")] + Windows936, + #[value(name = "windows-949")] + Windows949, + #[value(name = "windows-950")] + Windows950, + #[value(name = "windows-1250")] + Windows1250, + #[value(name = "windows-1251")] + Windows1251, + #[value(name = "windows-1252")] + Windows1252, + #[value(name = "windows-1253")] + Windows1253, + #[value(name = "windows-1254")] + Windows1254, + #[value(name = "windows-1255")] + Windows1255, + #[value(name = "windows-1256")] + Windows1256, + #[value(name = "windows-1257")] + Windows1257, + #[value(name = "windows-1258")] + Windows1258, +} + +impl From for HeaderCodepage { + fn from(value: HeaderCodepageArg) -> Self { + match value { + HeaderCodepageArg::Ascii => HeaderCodepage::Ascii, + HeaderCodepageArg::Windows874 => HeaderCodepage::Windows874, + HeaderCodepageArg::Windows932 => HeaderCodepage::Windows932, + HeaderCodepageArg::Windows936 => HeaderCodepage::Windows936, + HeaderCodepageArg::Windows949 => HeaderCodepage::Windows949, + HeaderCodepageArg::Windows950 => HeaderCodepage::Windows950, + HeaderCodepageArg::Windows1250 => HeaderCodepage::Windows1250, + HeaderCodepageArg::Windows1251 => HeaderCodepage::Windows1251, + HeaderCodepageArg::Windows1252 => HeaderCodepage::Windows1252, + HeaderCodepageArg::Windows1253 => HeaderCodepage::Windows1253, + HeaderCodepageArg::Windows1254 => HeaderCodepage::Windows1254, + HeaderCodepageArg::Windows1255 => HeaderCodepage::Windows1255, + HeaderCodepageArg::Windows1256 => HeaderCodepage::Windows1256, + HeaderCodepageArg::Windows1257 => HeaderCodepage::Windows1257, + HeaderCodepageArg::Windows1258 => HeaderCodepage::Windows1258, + } + } +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, ValueEnum)] +pub enum DateFormatArg { + #[value(name = "ctime")] + Ctime, + #[value(name = "dm")] + DayMonth, + #[value(name = "md")] + MonthDay, + #[value(name = "iso8601")] + Iso8601, +} + +impl From for EwfInfoDateFormat { + fn from(value: DateFormatArg) -> Self { + match value { + DateFormatArg::Ctime => EwfInfoDateFormat::Ctime, + DateFormatArg::DayMonth => EwfInfoDateFormat::DayMonth, + DateFormatArg::MonthDay => EwfInfoDateFormat::MonthDay, + DateFormatArg::Iso8601 => EwfInfoDateFormat::Iso8601, + } + } +} + +impl Cli { + pub fn run(self) -> miette::Result<()> { + if self.version { + println!("ewfinfo {}", env!("CARGO_PKG_VERSION")); + return Ok(()); + } + + let input = self + .ewf_files + .first() + .ok_or_else(|| miette!("missing EWF input path"))?; + + // Open bodyfile early (libewf fails early if the path cannot be opened). + let mut bodyfile_stream = if let Some(path) = &self.bodyfile { + Some( + std::fs::File::create(path) + .into_diagnostic() + .wrap_err_with(|| format!("unable to create bodyfile `{}`", path.display()))?, + ) + } else { + None + }; + + enum Mode<'a> { + FileEntry(&'a str), + Hierarchy, + Image, + } + + let mode = if let Some(query) = self.file_entry.as_deref() { + Mode::FileEntry(query) + } else if self.hierarchy { + Mode::Hierarchy + } else { + Mode::Image + }; + + match mode { + Mode::FileEntry(query) => { + if self.format == OutputFormat::Dfxml { + // libewf supports DFXML for these modes too, but we keep this binary-only + // surface area small while wiring up image-mode parity. + return Err(miette!( + "dfxml output is not yet implemented for -F/-H/-B logical-evidence modes" + )); + } + + // libewf prints a version header for text output in all modes. + println!("ewfinfo {}", env!("CARGO_PKG_VERSION")); + println!(); + + // Logical evidence operations require LEF inputs (`.L01`/`.Lx01`). + let lef = LefReader::open(input).map_err(|e| miette!("{e}"))?; + let entry = logical::find_entry_by_path(lef.entries(), query) + .ok_or_else(|| miette!("file entry not found: `{query}`"))?; + + if let Some(stream) = bodyfile_stream.as_mut() { + stream + .write_all( + bodyfile::render_bodyfile_line(entry, self.separator.as_char()) + .as_bytes(), + ) + .into_diagnostic() + .wrap_err("unable to write bodyfile line")?; + stream + .flush() + .into_diagnostic() + .wrap_err("unable to flush bodyfile")?; + return Ok(()); + } + + print!( + "{}", + logical::render_file_entry_text(entry, self.separator.as_char()) + ); + Ok(()) + } + Mode::Hierarchy => { + if self.format == OutputFormat::Dfxml { + return Err(miette!( + "dfxml output is not yet implemented for -F/-H/-B logical-evidence modes" + )); + } + + // libewf prints a version header for text output in all modes. + println!("ewfinfo {}", env!("CARGO_PKG_VERSION")); + println!(); + + let lef = LefReader::open(input).map_err(|e| miette!("{e}"))?; + + if let Some(stream) = bodyfile_stream.as_mut() { + stream + .write_all( + bodyfile::render_bodyfile(lef.entries(), self.separator.as_char()) + .as_bytes(), + ) + .into_diagnostic() + .wrap_err("unable to write bodyfile")?; + stream + .flush() + .into_diagnostic() + .wrap_err("unable to flush bodyfile")?; + return Ok(()); + } + + print!( + "{}", + logical::render_hierarchy_text(lef.entries(), self.separator.as_char()) + ); + Ok(()) + } + Mode::Image => { + let sections = if self.acquiry_only { + EwfInfoSections::AcquiryOnly + } else if self.media_only { + EwfInfoSections::MediaOnly + } else if self.errors_only { + EwfInfoSections::ErrorsOnly + } else { + EwfInfoSections::All + }; + + let print_options = EwfInfoPrintOptions { + date_format: self.date_format.into(), + sections, + color: EwfInfoColorMode::Auto, + }; + + let img = EwfReader::open(input).map_err(|e| miette!("{e}"))?; + let header_codepage: HeaderCodepage = self.header_codepage.into(); + if matches!(img.format(), EwfFormat::E01 | EwfFormat::S01) + && header_codepage != HeaderCodepage::Ascii + { + return Err(miette!( + "TODO: EWF1 header codepage `{}` decoding is not implemented yet", + header_codepage + )); + } + + let meta = img.image_metadata().map_err(|e| miette!("{e}"))?; + let report = + EwfInfoReport::from_image_metadata(&meta).map_err(|e| miette!("{e}"))?; + + match self.format { + OutputFormat::Text => { + // libewf prints a version header for text output. + println!("ewfinfo {}", env!("CARGO_PKG_VERSION")); + println!(); + let text = report.to_text(&print_options).map_err(|e| miette!("{e}"))?; + print!("{text}"); + Ok(()) + } + OutputFormat::Dfxml => { + let xml = report + .to_dfxml(&print_options) + .map_err(|e| miette!("{e}"))?; + print!("{xml}"); + Ok(()) + } + } + } + } + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_clap_parses_bodyfile_glued_short_opt() { + let cli = Cli::try_parse_from(["ewfinfo", "-Bbodyfile", "-H", "case.L01"]).unwrap(); + assert_eq!( + cli.bodyfile.as_deref(), + Some(std::path::Path::new("bodyfile")) + ); + assert!(cli.hierarchy); + } + + #[test] + fn test_clap_parses_separator_backslash() { + let cli = Cli::try_parse_from(["ewfinfo", "-H", "-s", "\\", "case.L01"]).unwrap(); + assert_eq!(cli.separator, PathSeparator::Backslash); + } + + #[test] + fn test_clap_rejects_conflicting_info_flags() { + let err = Cli::try_parse_from(["ewfinfo", "-e", "-i", "image.E01"]).unwrap_err(); + let msg = err.to_string(); + assert!(msg.contains("cannot be used with") || msg.contains("conflicts with")); + } +} diff --git a/crates/ewf/src/bin/ewfinfo/ewfinfo/mod.rs b/crates/ewf/src/bin/ewfinfo/ewfinfo/mod.rs new file mode 100644 index 0000000..63269b3 --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/ewfinfo/mod.rs @@ -0,0 +1,744 @@ +//! `ewfinfo`-style reporting for **image** (disk) EWF sets. +//! +//! This module provides a Rust-native, strongly-typed “report + rendering” API intended to back an +//! `ewfinfo`-compatible CLI. The CLI is expected to be responsible for argument parsing (clap), +//! selecting which sections to print (e.g. `-i`/`-m`/`-e`), and for **logical evidence** outputs +//! (`-F`/`-H`/`-B`). The `ewf` library only exposes spec-oriented metadata extraction; this module +//! owns the *image metadata* report model and renderers. +//! +//! ## Compatibility notes +//! +//! - The section structure and formatting are based on libewf’s `ewfinfo` implementation, but the +//! data model here is not a 1:1 port of libewf’s `info_handle_t`. Instead, the library exposes +//! stable Rust types that the application can map its CLI options into. +//! - DFXML output is **schema-aligned** (DFXML Working Group) and therefore intentionally differs +//! from libewf’s historic “DFXML” output (which uses an `ewfobjects` root). +//! - Any unsupported surface area must return an explicit [`EwfInfoError::Unsupported`] with a +//! clear `TODO:` marker rather than silently degrading behavior. +//! +//! ## References +//! +//! - `external/libewf/ewftools/info_handle.h` +//! - `external/libewf/ewftools/info_handle.c` +//! - `external/libewf/ewftools/ewfinfo.c` +//! - `external/libewf/manuals/ewfinfo.1` +//! - `external/refs/repos/dfxml-working-group__dfxml_schema.commit` +//! - `crates/dfxml/schema/dfxml.xsd` +//! +//! ## Where it is used +//! +//! The `ewfinfo` binary wires this module up in `crates/ewf/src/bin/ewfinfo/cli.rs`. + +mod print_dfxml; +mod print_text; + +use std::fmt; + +/// Error type for `ewfinfo` report building and printing. +/// +/// This error is intended for **library** usage, and uses `thiserror` for structured errors. +#[derive(Debug, thiserror::Error)] +pub enum EwfInfoError { + /// The underlying EWF reader returned an error. + #[error(transparent)] + EwF(#[from] ewf::Error), + + /// The operation is not supported yet. + /// + /// This variant is intentionally used for explicit, user-visible `TODO:` gaps when porting + /// libewf behavior. + #[error("unsupported: {0}")] + #[allow(dead_code)] + Unsupported(String), + + /// The report data is internally inconsistent (should generally be treated as a bug). + #[error("invalid ewfinfo report: {0}")] + InvalidReport(String), +} + +/// Result type used by this module. +pub type EwfInfoResult = std::result::Result; + +/// Header codepage for decoding EWF1 `header` section strings. +/// +/// libewf exposes this via `ewfinfo -A`; see `external/libewf/manuals/ewfinfo.1`. +#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)] +pub enum HeaderCodepage { + /// ASCII (libewf default). + #[default] + Ascii, + Windows874, + Windows932, + Windows936, + Windows949, + Windows950, + Windows1250, + Windows1251, + Windows1252, + Windows1253, + Windows1254, + Windows1255, + Windows1256, + Windows1257, + Windows1258, +} + +/// Date formatting mode used when printing timestamps. +/// +/// This corresponds to `ewfinfo -d` (`ctime`, `dm`, `md`, `iso8601`). +#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)] +pub enum EwfInfoDateFormat { + /// libewf default. + #[default] + Ctime, + /// Day/month (`dm`). + DayMonth, + /// Month/day (`md`). + MonthDay, + /// ISO-8601. + Iso8601, +} + +/// ANSI color mode for text output. +#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)] +pub enum EwfInfoColorMode { + /// Never emit ANSI escape sequences. + /// + /// This is the default to keep unit tests and golden outputs deterministic. + #[default] + Never, + /// Emit colors only when stdout is a terminal (and `NO_COLOR` is not set). + Auto, + /// Always emit colors (even when redirected). + #[allow(dead_code)] + Always, +} + +/// Options that affect how an [`EwfInfoReport`] is printed (formatting). +#[derive(Debug, Clone, Default)] +pub struct EwfInfoPrintOptions { + /// Date format (`ewfinfo -d`). + pub date_format: EwfInfoDateFormat, + + /// Which logical section set to render (maps to libewf `ewfinfo`’s `-i`/`-m`/`-e` flags). + pub sections: EwfInfoSections, + + /// ANSI color mode for text output. + pub color: EwfInfoColorMode, +} + +/// Selects which report sections to render. +#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)] +pub enum EwfInfoSections { + /// Render the full report (libewf `info_option == 'a'`). + #[default] + All, + /// Render only acquisition/header values (libewf `info_option == 'i'`). + AcquiryOnly, + /// Render only media-related information (libewf `info_option == 'm'`). + MediaOnly, + /// Render only acquisition read errors (libewf `info_option == 'e'`). + ErrorsOnly, +} + +/// A typed `ewfinfo` report for image metadata. +/// +/// The report is designed so that rendering can reproduce libewf `ewfinfo` output deterministically +/// (including section ordering). +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct EwfInfoReport { + /// Image filenames (segment set) used in DFXML output. + pub image_filenames: Vec, + + /// Human-oriented acquisition/header values (aka “Acquiry information”). + pub acquiry_information: Vec, + + /// EWF format information (“EWF information”). + pub ewf_information: Vec, + + /// Media geometry/flags (“Media information”). + pub media_information: Vec, + + /// Stored digest hashes (“Digest hash information”). + pub digest_hash_information: Vec, + + /// Session runs (may be empty). + pub sessions: Vec, + + /// Track runs (may be empty). + pub tracks: Vec, + + /// Acquisition read error runs (may be empty). + pub acquisition_read_errors: Vec, + + /// Bytes per sector used when printing runs. + pub bytes_per_sector: u32, +} + +impl EwfInfoReport { + /// Build a report from spec-oriented image metadata. + pub fn from_image_metadata(meta: &ewf::metadata::ImageMetadata) -> EwfInfoResult { + use ewf::metadata::{CompressionLevel, MediaType}; + use ewf::{EwfCompression, EwfFormat}; + + let image_filenames: Vec = meta + .segment_paths + .iter() + .map(|p| p.display().to_string()) + .collect(); + + let is_ewf1 = matches!(meta.format, EwfFormat::E01 | EwfFormat::S01); + + // Acquiry information is only available for EWF1 header sections. + let mut acquiry_information: Vec = Vec::new(); + if is_ewf1 { + push_string_field( + &mut acquiry_information, + "case_number", + "Case number", + meta.header_values.case_number.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "description", + "Description", + meta.header_values.description.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "examiner_name", + "Examiner name", + meta.header_values.examiner_name.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "evidence_number", + "Evidence number", + meta.header_values.evidence_number.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "notes", + "Notes", + meta.header_values.notes.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "acquiry_date", + "Acquisition date", + meta.header_values.acquisition_datetime.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "system_date", + "System date", + meta.header_values.system_datetime.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "acquiry_operating_system", + "Operating system used", + meta.header_values.acquisition_os.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "acquiry_software", + "Software used", + meta.header_values.acquisition_software.as_deref(), + ); + push_string_field( + &mut acquiry_information, + "acquiry_software_version", + "Software version used", + meta.header_values.acquisition_software_version.as_deref(), + ); + + // Password is special in libewf; we store the raw value (empty means “not set”). + acquiry_information.push(InfoField { + identifier: "password", + description: "Password", + value: InfoValue::String(meta.header_values.password.clone().unwrap_or_default()), + }); + } + + let file_format_str = meta.file_format.to_string(); + + let compression_method_str = match meta.compression_method { + EwfCompression::None => "none", + EwfCompression::Zlib => "deflate", + EwfCompression::Bzip2 => "bzip2", + EwfCompression::Unknown(_) => "unknown", + }; + + let compression_level_str = match meta.compression_level { + CompressionLevel::NoCompression => "no compression", + CompressionLevel::GoodFastCompression => "good (fast) compression", + CompressionLevel::BestCompression => "best compression", + CompressionLevel::Unknown => "unknown compression", + CompressionLevel::NotRecorded => "not recorded", + }; + + let mut ewf_information: Vec = Vec::new(); + ewf_information.push(InfoField { + identifier: "file_format", + description: "File format", + value: InfoValue::String(file_format_str), + }); + if let Some((maj, min)) = meta.segment_file_version { + ewf_information.push(InfoField { + identifier: "segment_file_version", + description: "Segment file version", + value: InfoValue::String(format!("{maj}.{min}")), + }); + } + ewf_information.push(InfoField { + identifier: "sectors_per_chunk", + description: "Sectors per chunk", + value: InfoValue::U32(meta.sectors_per_chunk), + }); + ewf_information.push(InfoField { + identifier: "error_granularity", + description: "Error granularity", + value: InfoValue::U32(meta.error_granularity), + }); + ewf_information.push(InfoField { + identifier: "compression_method", + description: "Compression method", + value: InfoValue::String(compression_method_str.to_string()), + }); + ewf_information.push(InfoField { + identifier: "compression_level", + description: "Compression level", + value: InfoValue::String(compression_level_str.to_string()), + }); + if let Some(set_id) = meta.set_identifier { + ewf_information.push(InfoField { + identifier: "set_identifier", + description: "Set identifier", + value: InfoValue::String(format_guid_le(&set_id)), + }); + } + + let media_type_str = match meta.media_type { + MediaType::RemovableDisk => "removable disk", + MediaType::FixedDisk => "fixed disk", + MediaType::OpticalDisk => "optical disk (CD/DVD/BD)", + MediaType::SingleFiles => "single files", + MediaType::MemoryRam => "memory (RAM)", + MediaType::Unknown => "unknown", + }; + + let media_information: Vec = vec![ + InfoField { + identifier: "media_type", + description: "Media type", + value: InfoValue::String(media_type_str.to_string()), + }, + InfoField { + identifier: "is_physical", + description: "Is physical", + value: InfoValue::Bool(meta.is_physical), + }, + InfoField { + identifier: "bytes_per_sector", + description: "Bytes per sector", + value: InfoValue::U32(meta.bytes_per_sector), + }, + InfoField { + identifier: "number_of_sectors", + description: "Number of sectors", + value: InfoValue::U64(meta.number_of_sectors), + }, + InfoField { + identifier: "media_size", + description: "Media size", + value: InfoValue::Size(meta.media_size), + }, + ]; + + let mut digest_hash_information: Vec = Vec::new(); + if let Some(md5) = meta.digests.md5 { + digest_hash_information.push(InfoField { + identifier: "md5", + description: "MD5", + value: InfoValue::String(hex_lower(&md5)), + }); + } + if let Some(sha1) = meta.digests.sha1 { + digest_hash_information.push(InfoField { + identifier: "sha1", + description: "SHA1", + value: InfoValue::String(hex_lower(&sha1)), + }); + } + + Ok(EwfInfoReport { + image_filenames, + acquiry_information, + ewf_information, + media_information, + digest_hash_information, + sessions: meta.sessions.clone(), + tracks: meta.tracks.clone(), + acquisition_read_errors: meta.acquisition_read_errors.clone(), + bytes_per_sector: meta.bytes_per_sector, + }) + } + + /// Render the report as a human-friendly text report. + /// + /// Unlike the binary formats, the text output is not intended to be 1:1 compatible with + /// libewf’s historical `ewfinfo` formatting. + pub fn to_text(&self, options: &EwfInfoPrintOptions) -> EwfInfoResult { + print_text::render_text(self, options) + } + + /// Render the report as schema-aligned DFXML (DFXML 2.0.0-beta.0). + pub fn to_dfxml(&self, options: &EwfInfoPrintOptions) -> EwfInfoResult { + print_dfxml::render_dfxml(self, options) + } +} + +/// A single field/value printed inside a section. +/// +/// The `identifier` is used for DFXML element naming (it matches libewf’s identifiers), while the +/// `description` is used for text output labels. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct InfoField { + /// DFXML identifier (e.g. `case_number`, `bytes_per_sector`). + pub identifier: &'static str, + /// Text label (e.g. `Case number`, `Bytes per sector`). + pub description: &'static str, + /// Field value (typed). + pub value: InfoValue, +} + +/// Field value variants needed to reproduce libewf formatting. +#[derive(Debug, Clone, PartialEq, Eq)] +pub enum InfoValue { + /// Arbitrary string value. + String(String), + /// Unsigned 32-bit integer. + U32(u32), + /// Unsigned 64-bit integer. + U64(u64), + /// Size value in bytes (text output prints human-readable MiB + bytes). + Size(u64), + /// Boolean (“yes”/“no”). + Bool(bool), +} + +fn push_string_field( + dst: &mut Vec, + identifier: &'static str, + description: &'static str, + value: Option<&str>, +) { + let Some(value) = value else { + return; + }; + if value.is_empty() { + return; + } + dst.push(InfoField { + identifier, + description, + value: InfoValue::String(value.to_string()), + }); +} + +fn hex_lower(bytes: &[u8]) -> String { + const TABLE: &[u8; 16] = b"0123456789abcdef"; + let mut out = String::with_capacity(bytes.len() * 2); + for b in bytes { + out.push(TABLE[(b >> 4) as usize] as char); + out.push(TABLE[(b & 0x0f) as usize] as char); + } + out +} + +fn format_guid_le(guid: &[u8; 16]) -> String { + // Matches libewf’s little-endian GUID formatting used by `ewfinfo` for set identifiers. + format!( + "{:02x}{:02x}{:02x}{:02x}-{:02x}{:02x}-{:02x}{:02x}-{:02x}{:02x}-{:02x}{:02x}{:02x}{:02x}{:02x}{:02x}", + guid[3], + guid[2], + guid[1], + guid[0], + guid[5], + guid[4], + guid[7], + guid[6], + guid[8], + guid[9], + guid[10], + guid[11], + guid[12], + guid[13], + guid[14], + guid[15] + ) +} + +fn ensure_bytes_per_sector(report: &EwfInfoReport) -> EwfInfoResult { + if report.bytes_per_sector == 0 { + return Err(EwfInfoError::InvalidReport( + "bytes_per_sector must be non-zero".to_string(), + )); + } + Ok(report.bytes_per_sector) +} + +fn format_bool_yes_no(v: bool) -> &'static str { + if v { "yes" } else { "no" } +} + +fn format_size_mib(bytes: u64) -> String { + // Mirrors libewf’s `byte_size_string_create(..., MEBIBYTE)` behavior: choose a MiB unit. + // We keep it simple (no locale), and fall back to raw bytes where formatting is ambiguous. + // + // NOTE: This is intentionally deterministic; we do not attempt “smart” unit switching. + let mib = 1024u64 * 1024u64; + if bytes < mib { + return format!("{bytes} bytes"); + } + let value = (bytes as f64) / (mib as f64); + // libewf prints e.g. "1.4 MiB" for 1474560. That is one decimal for non-integer values. + let s = if (value - value.round()).abs() < f64::EPSILON { + format!("{:.0} MiB", value) + } else { + format!("{:.1} MiB", value) + }; + format!("{s} ({bytes} bytes)") +} + +fn format_datetime_value(value: &str, fmt: EwfInfoDateFormat) -> String { + // libewf supports multiple date formats for header values (`ewfinfo -d`). + // + // Empirically (via libewf’s `ewfinfo`), the input header values are commonly encoded as: + // + // - `YYYY-MM-DD HH:MM:SS` + // - Unix epoch seconds (e.g. `1361530430`) + // + // and then rendered in one of the following formats: + // + // - `ctime`: `Fri Feb 22 12:53:50 2013` + // - `iso8601`: `2013-02-22T12:53:50` + // - `dm`: `22/02/2013 12:53:50` + // - `md`: `02/22/2013 12:53:50` + // + // We use `jiff`'s `strptime`/`strftime`-style routines to avoid hand-rolled parsing and to + // correctly compute weekday names for `ctime`. + // + // If parsing fails, we fall back to the original string (no silent coercion). + let value = value.trim(); + + // 1) Common case: `YYYY-MM-DD HH:MM:SS` + if let Ok(dt) = jiff::fmt::strtime::parse("%F %T", value).and_then(|tm| tm.to_datetime()) { + return match fmt { + EwfInfoDateFormat::Ctime => dt.strftime("%a %b %e %T %Y").to_string(), + EwfInfoDateFormat::Iso8601 => dt.strftime("%FT%T").to_string(), + EwfInfoDateFormat::DayMonth => dt.strftime("%d/%m/%Y %T").to_string(), + EwfInfoDateFormat::MonthDay => dt.strftime("%m/%d/%Y %T").to_string(), + }; + } + + // 2) Unix epoch seconds (`time_t` style). libewf renders these in the system time zone. + // We accept an optional fractional component but ignore it for display (libewf prints seconds). + let secs_str = value.split_once('.').map(|(s, _)| s).unwrap_or(value); + if let Ok(secs) = secs_str.parse::() { + let tz = jiff::tz::TimeZone::try_system().unwrap_or(jiff::tz::TimeZone::UTC); + if let Ok(ts) = jiff::Timestamp::from_second(secs) { + let dt = ts.to_zoned(tz).datetime(); + return match fmt { + EwfInfoDateFormat::Ctime => dt.strftime("%a %b %e %T %Y").to_string(), + EwfInfoDateFormat::Iso8601 => dt.strftime("%FT%T").to_string(), + EwfInfoDateFormat::DayMonth => dt.strftime("%d/%m/%Y %T").to_string(), + EwfInfoDateFormat::MonthDay => dt.strftime("%m/%d/%Y %T").to_string(), + }; + } + } + + value.to_string() +} + +impl fmt::Display for HeaderCodepage { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + let s = match self { + HeaderCodepage::Ascii => "ascii", + HeaderCodepage::Windows874 => "windows-874", + HeaderCodepage::Windows932 => "windows-932", + HeaderCodepage::Windows936 => "windows-936", + HeaderCodepage::Windows949 => "windows-949", + HeaderCodepage::Windows950 => "windows-950", + HeaderCodepage::Windows1250 => "windows-1250", + HeaderCodepage::Windows1251 => "windows-1251", + HeaderCodepage::Windows1252 => "windows-1252", + HeaderCodepage::Windows1253 => "windows-1253", + HeaderCodepage::Windows1254 => "windows-1254", + HeaderCodepage::Windows1255 => "windows-1255", + HeaderCodepage::Windows1256 => "windows-1256", + HeaderCodepage::Windows1257 => "windows-1257", + HeaderCodepage::Windows1258 => "windows-1258", + }; + f.write_str(s) + } +} + +impl fmt::Display for EwfInfoDateFormat { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + let s = match self { + EwfInfoDateFormat::Ctime => "ctime", + EwfInfoDateFormat::DayMonth => "dm", + EwfInfoDateFormat::MonthDay => "md", + EwfInfoDateFormat::Iso8601 => "iso8601", + }; + f.write_str(s) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + fn sample_report(bytes_per_sector: u32) -> EwfInfoReport { + EwfInfoReport { + image_filenames: vec!["image.E01".to_string()], + acquiry_information: vec![InfoField { + identifier: "case_number", + description: "Case number", + value: InfoValue::String("1".to_string()), + }], + ewf_information: vec![InfoField { + identifier: "sectors_per_chunk", + description: "Sectors per chunk", + value: InfoValue::U32(64), + }], + media_information: vec![ + InfoField { + identifier: "bytes_per_sector", + description: "Bytes per sector", + value: InfoValue::U32(bytes_per_sector), + }, + InfoField { + identifier: "media_size", + description: "Media size", + value: InfoValue::Size(1_474_560), + }, + ], + digest_hash_information: vec![InfoField { + identifier: "md5", + description: "MD5", + value: InfoValue::String("ae1ce8f5ac079d3ee93f97fe3792bda3".to_string()), + }], + sessions: vec![ewf::metadata::SectorRun { + start_sector: 0, + sector_count: 10, + }], + tracks: vec![], + acquisition_read_errors: vec![ewf::metadata::SectorRun { + start_sector: 100, + sector_count: 0, + }], + bytes_per_sector, + } + } + + #[test] + fn test_text_rendering_layout_and_sections() { + let report = sample_report(512); + let out = report.to_text(&EwfInfoPrintOptions::default()).unwrap(); + + let expected = concat!( + "Acquiry information\n", + "───────────────────\n", + " Case number: 1\n", + "\n", + "EWF information\n", + "───────────────\n", + " Sectors per chunk: 64\n", + "\n", + "Media information\n", + "─────────────────\n", + " Bytes per sector: 512\n", + " Media size: 1.4 MiB (1474560 bytes)\n", + "\n", + "Digest hash information\n", + "───────────────────────\n", + " MD5: ae1ce8f5ac079d3ee93f97fe3792bda3\n", + "\n", + "Sessions (1)\n", + "────────────\n", + " - sectors 0..9 (10 sectors)\n", + "\n", + "Read errors during acquisition (1)\n", + "──────────────────────────────────\n", + " - sectors 100..100 (0 sectors)\n", + "\n", + ); + + assert_eq!(out, expected); + } + + #[test] + fn test_dfxml_rendering_emits_runs_and_escapes() { + let mut report = sample_report(512); + report.acquiry_information.push(InfoField { + identifier: "acquiry_date", + description: "Acquisition date", + value: InfoValue::String("2020-01-01 & ".to_string()), + }); + + let out = report.to_dfxml(&EwfInfoPrintOptions::default()).unwrap(); + + assert!(out.contains("")); + assert!(out.contains("Disk Image")); + + // Acquiry fields are mapped into Dublin Core for schema-aligned DFXML. + assert!(out.contains("1")); + assert!(out.contains("2020-01-01 & <test>")); + + // Runs are printed in bytes. + assert!(out.contains("img_offset=\"0\"")); + assert!(out.contains("len=\"5120\"")); + // sector_count=0 yields len=0; start sector 100 -> 100*512=51200 + assert!(out.contains("img_offset=\"51200\"")); + assert!(out.contains("len=\"0\"")); + } + + #[test] + fn test_renderers_respect_section_filtering() { + let report = sample_report(512); + + let acquiry_only = EwfInfoPrintOptions { + sections: EwfInfoSections::AcquiryOnly, + ..EwfInfoPrintOptions::default() + }; + let out = report.to_text(&acquiry_only).unwrap(); + assert!(out.contains("Acquiry information\n")); + assert!(!out.contains("EWF information\n")); + assert!(!out.contains("Media information\n")); + assert!(!out.contains("Read errors during acquisition")); + + let errors_only = EwfInfoPrintOptions { + sections: EwfInfoSections::ErrorsOnly, + ..EwfInfoPrintOptions::default() + }; + let out = report.to_dfxml(&errors_only).unwrap(); + // Metadata is always present, but acquisition-only Dublin Core entries should be omitted. + assert!(!out.contains("")); + // Media-only image hash run should be omitted. + assert!(!out.contains("type=\"image\"")); + // Error byte_runs should still render. + assert!(out.contains("type=\"acquisition_read_error\"")); + } + + #[test] + fn test_renderers_require_bytes_per_sector() { + let report = sample_report(0); + let err = report.to_text(&EwfInfoPrintOptions::default()).unwrap_err(); + assert!(matches!(err, EwfInfoError::InvalidReport(_))); + } +} diff --git a/crates/ewf/src/bin/ewfinfo/ewfinfo/print_dfxml.rs b/crates/ewf/src/bin/ewfinfo/ewfinfo/print_dfxml.rs new file mode 100644 index 0000000..130b23b --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/ewfinfo/print_dfxml.rs @@ -0,0 +1,196 @@ +use super::{ + EwfInfoPrintOptions, EwfInfoReport, EwfInfoResult, EwfInfoSections, InfoField, InfoValue, + ensure_bytes_per_sector, format_datetime_value, +}; + +use dfxml::{ + BuildEnvironment, ByteRun, Creator, DfxmlDocument, DiskImageObject, DublinCoreElement, + DublinCoreElementName, HashDigest, HashDigestType, Library, Source, +}; + +/// Render [`EwfInfoReport`] as schema-aligned DFXML (DFXML 2.0.0-beta.0). +/// +/// Unlike libewf’s historic `ewfinfo -f dfxml` output (which uses an `ewfobjects` root), this +/// emits a schema-aligned `` document. +pub(super) fn render_dfxml( + report: &EwfInfoReport, + options: &EwfInfoPrintOptions, +) -> EwfInfoResult { + let bps = ensure_bytes_per_sector(report)? as u64; + + let want_acquiry = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::AcquiryOnly + ); + let want_media = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::MediaOnly + ); + let want_errors = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::ErrorsOnly + ); + + let mut doc = DfxmlDocument::new(); + + // Metadata is required by schema. + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Type, + value: "Disk Image".to_string(), + }); + + if want_acquiry { + if let Some(s) = get_string(&report.acquiry_information, "description") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Title, + value: s.to_string(), + }); + } + if let Some(s) = get_string(&report.acquiry_information, "examiner_name") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Creator, + value: s.to_string(), + }); + } + if let Some(s) = get_string(&report.acquiry_information, "case_number") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Identifier, + value: s.to_string(), + }); + } + if let Some(s) = get_string(&report.acquiry_information, "evidence_number") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Identifier, + value: s.to_string(), + }); + } + if let Some(s) = get_string(&report.acquiry_information, "notes") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Description, + value: s.to_string(), + }); + } + if let Some(s) = get_string(&report.acquiry_information, "acquiry_date") { + doc.metadata.dublin_core.push(DublinCoreElement { + name: DublinCoreElementName::Date, + value: format_datetime_value(s, options.date_format), + }); + } + } + + doc.creator = Some(Creator { + program: Some("ewfinfo".to_string()), + version: Some(env!("CARGO_PKG_VERSION").to_string()), + build_environment: Some(BuildEnvironment { + compiler: Some("rustc".to_string()), + compilation_date: None, + libraries: vec![Library { + name: Some("ewf".to_string()), + version: Some(env!("CARGO_PKG_VERSION").to_string()), + }], + }), + execution_environment: Some(dfxml::ExecutionEnvironment { + os_sysname: Some(std::env::consts::OS.to_string()), + arch: Some(std::env::consts::ARCH.to_string()), + }), + }); + + if !report.image_filenames.is_empty() { + doc.source = Some(Source { + image_filenames: report.image_filenames.clone(), + }); + } + + if want_media || want_errors { + let mut dio = DiskImageObject { + sector_size: Some(bps), + byte_runs: Vec::new(), + }; + + if want_media { + // Map image-level hashes to a single byte_run over the whole image (when possible). + if let Some(media_size) = get_media_size_bytes(report) { + let mut image_run = ByteRun { + img_offset: 0, + len: media_size, + kind: Some("image".to_string()), + hashdigests: Vec::new(), + }; + for h in &report.digest_hash_information { + let InfoValue::String(s) = &h.value else { + continue; + }; + let algorithm = match h.identifier { + "md5" => HashDigestType::Md5, + "sha1" => HashDigestType::Sha1, + _ => continue, + }; + image_run.hashdigests.push(HashDigest { + algorithm, + value: s.clone(), + }); + } + dio.byte_runs.push(image_run); + } + + for r in &report.sessions { + dio.byte_runs.push(ByteRun { + img_offset: r.start_sector.saturating_mul(bps), + len: r.sector_count.saturating_mul(bps), + kind: Some("session".to_string()), + hashdigests: Vec::new(), + }); + } + for r in &report.tracks { + dio.byte_runs.push(ByteRun { + img_offset: r.start_sector.saturating_mul(bps), + len: r.sector_count.saturating_mul(bps), + kind: Some("track".to_string()), + hashdigests: Vec::new(), + }); + } + } + + if want_errors { + for r in &report.acquisition_read_errors { + dio.byte_runs.push(ByteRun { + img_offset: r.start_sector.saturating_mul(bps), + len: r.sector_count.saturating_mul(bps), + kind: Some("acquisition_read_error".to_string()), + hashdigests: Vec::new(), + }); + } + } + + doc.diskimageobjects.push(dio); + } + + doc.to_xml_string() + .map_err(|e| super::EwfInfoError::InvalidReport(format!("dfxml write failed: {e}"))) +} + +fn get_string<'a>(fields: &'a [InfoField], id: &str) -> Option<&'a str> { + fields.iter().find_map(|f| { + if f.identifier != id { + return None; + } + match &f.value { + InfoValue::String(s) if !s.is_empty() => Some(s.as_str()), + _ => None, + } + }) +} + +fn get_media_size_bytes(report: &EwfInfoReport) -> Option { + report.media_information.iter().find_map(|f| { + if f.identifier != "media_size" { + return None; + } + match f.value { + InfoValue::Size(v) => Some(v), + InfoValue::U64(v) => Some(v), + InfoValue::U32(v) => Some(v as u64), + _ => None, + } + }) +} diff --git a/crates/ewf/src/bin/ewfinfo/ewfinfo/print_text.rs b/crates/ewf/src/bin/ewfinfo/ewfinfo/print_text.rs new file mode 100644 index 0000000..601488b --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/ewfinfo/print_text.rs @@ -0,0 +1,306 @@ +use super::{ + EwfInfoColorMode, EwfInfoPrintOptions, EwfInfoReport, EwfInfoResult, EwfInfoSections, + InfoField, InfoValue, ensure_bytes_per_sector, format_bool_yes_no, format_datetime_value, + format_size_mib, +}; + +use std::io::IsTerminal as _; + +use textwrap::core::display_width; + +/// Render [`EwfInfoReport`] as a human-friendly text report. +pub(super) fn render_text( + report: &EwfInfoReport, + options: &EwfInfoPrintOptions, +) -> EwfInfoResult { + let _bps = ensure_bytes_per_sector(report)?; + let mut out = String::new(); + + let use_color = color_enabled(options.color); + let width = detect_width(options.color); + + let want_acquiry = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::AcquiryOnly + ); + let want_media = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::MediaOnly + ); + let want_errors = matches!( + options.sections, + EwfInfoSections::All | EwfInfoSections::ErrorsOnly + ); + + // Section: Acquiry information + if want_acquiry { + push_heading(&mut out, "Acquiry information", use_color); + if report.acquiry_information.is_empty() { + out.push_str(" (no information found)\n\n"); + } else { + let label_width = max_label_width(&report.acquiry_information).min(32); + for f in &report.acquiry_information { + let value = match (f.identifier, &f.value) { + ("password", InfoValue::String(s)) => { + if s.is_empty() { + "N/A".to_string() + } else { + format!("(hash: {s})") + } + } + ("acquiry_date", InfoValue::String(s)) + | ("system_date", InfoValue::String(s)) => { + format_datetime_value(s, options.date_format) + } + (_, InfoValue::String(s)) => s.clone(), + (_, InfoValue::U32(v)) => v.to_string(), + (_, InfoValue::U64(v)) => v.to_string(), + (_, InfoValue::Size(v)) => format_size_mib(*v), + (_, InfoValue::Bool(v)) => format_bool_yes_no(*v).to_string(), + }; + push_kv( + &mut out, + f.description, + &value, + label_width, + width, + use_color, + ); + } + out.push('\n'); + } + } + + if want_media { + // Section: EWF information + push_heading(&mut out, "EWF information", use_color); + let label_width = max_label_width(&report.ewf_information).min(32); + for f in &report.ewf_information { + let value = match &f.value { + InfoValue::String(s) => s.clone(), + InfoValue::U32(v) => v.to_string(), + InfoValue::U64(v) => v.to_string(), + InfoValue::Size(v) => format_size_mib(*v), + InfoValue::Bool(v) => format_bool_yes_no(*v).to_string(), + }; + push_kv( + &mut out, + f.description, + &value, + label_width, + width, + use_color, + ); + } + out.push('\n'); + + // Section: Media information + push_heading(&mut out, "Media information", use_color); + let label_width = max_label_width(&report.media_information).min(32); + for f in &report.media_information { + let value = match &f.value { + InfoValue::String(s) => s.clone(), + InfoValue::U32(v) => v.to_string(), + InfoValue::U64(v) => v.to_string(), + InfoValue::Size(v) => format_size_mib(*v), + InfoValue::Bool(v) => format_bool_yes_no(*v).to_string(), + }; + push_kv( + &mut out, + f.description, + &value, + label_width, + width, + use_color, + ); + } + out.push('\n'); + + // Section: Digest hash information (only if present). + if !report.digest_hash_information.is_empty() { + push_heading(&mut out, "Digest hash information", use_color); + let label_width = max_label_width(&report.digest_hash_information).min(32); + for f in &report.digest_hash_information { + let value = match &f.value { + InfoValue::String(s) => s.clone(), + InfoValue::U32(v) => v.to_string(), + InfoValue::U64(v) => v.to_string(), + InfoValue::Size(v) => format_size_mib(*v), + InfoValue::Bool(v) => format_bool_yes_no(*v).to_string(), + }; + push_kv( + &mut out, + f.description, + &value, + label_width, + width, + use_color, + ); + } + out.push('\n'); + } + + // Sessions / tracks: only print if present. + push_runs(&mut out, "Sessions", &report.sessions, width, use_color); + push_runs(&mut out, "Tracks", &report.tracks, width, use_color); + } + + if want_errors && !report.acquisition_read_errors.is_empty() { + push_runs( + &mut out, + "Read errors during acquisition", + &report.acquisition_read_errors, + width, + use_color, + ); + } + + Ok(out) +} + +fn max_label_width(fields: &[InfoField]) -> usize { + fields + .iter() + .map(|f| display_width(f.description)) + .max() + .unwrap_or(0) +} + +fn detect_width(color_mode: EwfInfoColorMode) -> usize { + // Keep deterministic output in non-interactive contexts. + if matches!(color_mode, EwfInfoColorMode::Always) || std::io::stdout().is_terminal() { + if let Ok(cols) = std::env::var("COLUMNS") + && let Ok(n) = cols.parse::() + && (40..=240).contains(&n) + { + return n; + } + 100 + } else { + 80 + } +} + +fn color_enabled(mode: EwfInfoColorMode) -> bool { + match mode { + EwfInfoColorMode::Never => false, + EwfInfoColorMode::Always => std::env::var_os("NO_COLOR").is_none(), + EwfInfoColorMode::Auto => { + std::io::stdout().is_terminal() && std::env::var_os("NO_COLOR").is_none() + } + } +} + +fn push_heading(out: &mut String, title: &str, color: bool) { + if color { + out.push_str("\x1b[1;36m"); + out.push_str(title); + out.push_str("\x1b[0m\n"); + + out.push_str("\x1b[2m"); + out.extend(std::iter::repeat_n('─', display_width(title))); + out.push_str("\x1b[0m\n"); + return; + } + + out.push_str(title); + out.push('\n'); + out.extend(std::iter::repeat_n('─', display_width(title))); + out.push('\n'); +} + +fn push_kv( + out: &mut String, + label: &str, + value: &str, + label_width: usize, + width: usize, + color: bool, +) { + const INDENT: &str = " "; + + let label_w = display_width(label); + let value_start = label_width.saturating_add(2); + let pad = value_start.saturating_sub(label_w.saturating_add(1)).max(1); + + let available = width + .saturating_sub(display_width(INDENT)) + .saturating_sub(value_start) + .max(1); + + let mut first_line = true; + for (para_i, para) in value.split('\n').enumerate() { + if para_i > 0 { + // Preserve explicit newlines. + out.push('\n'); + } + + let wrapped = textwrap::wrap(para, available); + for line in wrapped { + if first_line { + out.push_str(INDENT); + if color { + out.push_str("\x1b[2m"); + } + out.push_str(label); + out.push(':'); + out.extend(std::iter::repeat_n(' ', pad)); + if color { + out.push_str("\x1b[0m"); + } + out.push_str(&line); + out.push('\n'); + first_line = false; + } else { + out.push_str(INDENT); + out.extend(std::iter::repeat_n(' ', value_start)); + out.push_str(&line); + out.push('\n'); + } + } + } +} + +fn push_runs( + out: &mut String, + title: &str, + runs: &[ewf::metadata::SectorRun], + width: usize, + color: bool, +) { + if runs.is_empty() { + return; + } + + let heading = format!("{title} ({})", runs.len()); + push_heading(out, &heading, color); + + const INDENT: &str = " "; + const BULLET: &str = "- "; + let prefix = format!("{INDENT}{BULLET}"); + let prefix_w = display_width(&prefix); + let available = width.saturating_sub(prefix_w).max(1); + let subsequent = format!("{INDENT} "); + + for r in runs { + let mut last_sector = r.start_sector.saturating_add(r.sector_count); + if r.sector_count != 0 { + last_sector = last_sector.saturating_sub(1); + } + let msg = format!( + "sectors {}..{} ({} sectors)", + r.start_sector, last_sector, r.sector_count + ); + let wrapped = textwrap::wrap(&msg, available); + for (i, line) in wrapped.iter().enumerate() { + if i == 0 { + out.push_str(&prefix); + } else { + out.push_str(&subsequent); + } + out.push_str(line); + out.push('\n'); + } + } + out.push('\n'); +} diff --git a/crates/ewf/src/bin/ewfinfo/logical.rs b/crates/ewf/src/bin/ewfinfo/logical.rs new file mode 100644 index 0000000..b027c30 --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/logical.rs @@ -0,0 +1,141 @@ +//! Formatting helpers for `ewfinfo` logical-evidence modes (`-F` / `-H`). +//! +//! These outputs intentionally live in the binary crate (not `ewf`’s library API), mirroring +//! libewf’s `info_handle_file_entry_*` and `info_handle_logical_files_hierarchy_*` printing +//! functions. +//! +//! References: +//! - `external/libewf/ewftools/info_handle.c` +//! - `external/libewf/manuals/ewfinfo.1` + +use ewf::LefEntry; + +pub fn normalize_query_path(path: &str) -> String { + // Match `ewf::LefReader`’s normalization (`normalize_lef_path`). + path.replace('\\', "/").trim_start_matches("./").to_string() +} + +pub fn display_path(path: &str, separator: char) -> String { + match separator { + '/' => path.to_string(), + '\\' => path.replace('/', "\\"), + other => path.replace('/', &other.to_string()), + } +} + +pub fn find_entry_by_path<'a>(entries: &'a [LefEntry], query: &str) -> Option<&'a LefEntry> { + let want = normalize_query_path(query); + entries.iter().find(|e| e.path == want) +} + +pub fn render_hierarchy_text(entries: &[LefEntry], separator: char) -> String { + let mut out = String::new(); + for e in entries { + out.push_str(&display_path(&e.path, separator)); + out.push('\n'); + } + out +} + +pub fn render_file_entry_text(entry: &LefEntry, separator: char) -> String { + let mut out = String::new(); + + out.push_str("File entry information:\n"); + out.push_str(&format!( + "\tName\t\t\t\t: {}\n", + display_path(&entry.path, separator) + )); + out.push_str(&format!( + "\tType\t\t\t\t: {}\n", + if entry.is_dir { "directory" } else { "file" } + )); + + match entry.file_identifier { + Some(v) => out.push_str(&format!("\tFile identifier\t\t\t: {v}\n")), + None => out.push_str("\tFile identifier\t\t\t: N/A\n"), + } + + out.push_str(&format!("\tSize\t\t\t\t: {}\n", entry.size)); + + let fmt_time = |t: Option| match t { + Some(v) => v.to_string(), + None => "N/A".to_string(), + }; + + out.push_str(&format!( + "\tAccess time\t\t\t: {}\n", + fmt_time(entry.access_time) + )); + out.push_str(&format!( + "\tModification time\t\t: {}\n", + fmt_time(entry.modification_time) + )); + out.push_str(&format!( + "\tEntry modification time\t\t: {}\n", + fmt_time(entry.entry_modification_time) + )); + out.push_str(&format!( + "\tCreation time\t\t\t: {}\n", + fmt_time(entry.creation_time) + )); + + out.push('\n'); + + if entry.extents.is_empty() { + out.push_str("Extents: none\n"); + return out; + } + + out.push_str("Extents:\n"); + for (idx, ext) in entry.extents.iter().enumerate() { + out.push_str(&format!( + "\t{idx}\t\t\t\t: offset={} size={}\n", + ext.offset, ext.size + )); + } + out +} + +#[cfg(test)] +mod tests { + use super::*; + use ewf::LefExtent; + + #[test] + fn test_normalize_query_path_backslashes_and_dot_slash() { + assert_eq!(normalize_query_path(r".\dir\file.txt"), "dir/file.txt"); + } + + #[test] + fn test_render_hierarchy_text_respects_separator() { + let entries = vec![ + LefEntry { + path: "dir".to_string(), + is_dir: true, + size: 0, + extents: vec![], + file_identifier: Some(1), + access_time: None, + modification_time: None, + entry_modification_time: None, + creation_time: None, + }, + LefEntry { + path: "dir/file.txt".to_string(), + is_dir: false, + size: 5, + extents: vec![LefExtent { offset: 0, size: 5 }], + file_identifier: Some(2), + access_time: None, + modification_time: None, + entry_modification_time: None, + creation_time: None, + }, + ]; + + assert_eq!( + render_hierarchy_text(&entries, '\\'), + "dir\ndir\\file.txt\n" + ); + } +} diff --git a/crates/ewf/src/bin/ewfinfo/main.rs b/crates/ewf/src/bin/ewfinfo/main.rs new file mode 100644 index 0000000..522323c --- /dev/null +++ b/crates/ewf/src/bin/ewfinfo/main.rs @@ -0,0 +1,26 @@ +//! `ewfinfo` – show meta data stored in EWF files. +//! +//! This binary is the CLI front-end for the `ewf` crate. +//! +//! - Image metadata reporting (`-f text|dfxml`, `-i`/`-m`/`-e`, `-A`, `-d`) is implemented here +//! (binary-owned reporting/rendering), backed by spec-oriented metadata from `ewf::EwfReader`. +//! - Logical evidence outputs (`-F`, `-H`, `-B`) are **CLI-only** and intentionally live here (not +//! in the library), mirroring libewf’s separation between `info_handle` printing routines and the +//! core readers. +//! +//! References: +//! - `external/libewf/ewftools/ewfinfo.c` +//! - `external/libewf/ewftools/info_handle.c` +//! - `external/libewf/ewftools/bodyfile.c` +//! - `external/libewf/manuals/ewfinfo.1` + +mod bodyfile; +mod cli; +mod ewfinfo; +mod logical; + +use clap::Parser; + +fn main() -> miette::Result<()> { + cli::Cli::parse().run() +} diff --git a/crates/ewf/src/ewf1/file_header.rs b/crates/ewf/src/ewf1/file_header.rs new file mode 100644 index 0000000..8f295eb --- /dev/null +++ b/crates/ewf/src/ewf1/file_header.rs @@ -0,0 +1,117 @@ +//! EWF1 segment file header (E01/S01/L01). +//! +//! Every EWF1 segment starts with a small fixed-size header: +//! +//! - Offset `0x00..0x08` (`[u8; 8]`): signature (`EVF\t\r\n\xff\x00` or `LVF\t\r\n\xff\x00`) +//! - Offset `0x08` (`u8`): start-of-fields marker (`0x01`) +//! - Offset `0x09..0x0b` (`u16` LE): segment number +//! - Offset `0x0b..0x0d` (`u16` LE): end-of-fields marker (`0x0000`) +//! +//! References: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - libewf implementation: +//! - `external/libewf/libewf/libewf_file_header.c` +//! +//! This module uses `binrw` to keep the on-disk layout declarative and symmetric (read+write). + +use crate::{Error, Result}; +use binrw::{BinRead as _, BinWrite as _, binrw}; +use std::io::Cursor; + +use super::{EWF1_EVF_SIGNATURE, EWF1_FILE_HEADER_SIZE, EWF1_LVF_SIGNATURE}; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf1Signature { + Evf, + Lvf, +} + +impl Ewf1Signature { + pub(crate) fn bytes(self) -> [u8; 8] { + match self { + Self::Evf => EWF1_EVF_SIGNATURE, + Self::Lvf => EWF1_LVF_SIGNATURE, + } + } + + pub(crate) fn from_bytes(signature: [u8; 8]) -> Result { + if signature == EWF1_EVF_SIGNATURE { + Ok(Self::Evf) + } else if signature == EWF1_LVF_SIGNATURE { + Ok(Self::Lvf) + } else { + Err(Error::Invalid("unsupported EWF1 signature".to_string())) + } + } +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) struct Ewf1FileHeader { + #[br(try_map = Ewf1Signature::from_bytes)] + #[bw(map = |sig| sig.bytes())] + pub(crate) signature: Ewf1Signature, + #[br(assert(start_of_fields == 0x01))] + start_of_fields: u8, + pub(crate) segment_number: u16, + #[br(assert(end_of_fields == 0))] + end_of_fields: u16, +} + +impl Ewf1FileHeader { + pub(crate) fn new(signature: Ewf1Signature, segment_number: u16) -> Self { + Self { + signature, + start_of_fields: 0x01, + segment_number, + end_of_fields: 0, + } + } + + #[allow(dead_code)] + pub(crate) fn parse(bytes: &[u8; EWF1_FILE_HEADER_SIZE]) -> Result { + let mut cur = Cursor::new(bytes.as_slice()); + let hdr = Self::read(&mut cur) + .map_err(|e| Error::Invalid(format!("invalid EWF1 segment file header: {e}")))?; + Ok(hdr) + } + + pub(crate) fn to_bytes(self) -> [u8; EWF1_FILE_HEADER_SIZE] { + let mut out = [0u8; EWF1_FILE_HEADER_SIZE]; + self.write(&mut Cursor::new(&mut out[..])) + .expect("in-memory write cannot fail"); + out + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_roundtrip_evf_bytes() { + let hdr = Ewf1FileHeader::new(Ewf1Signature::Evf, 7); + let bytes = hdr.to_bytes(); + let parsed = Ewf1FileHeader::parse(&bytes).unwrap(); + assert_eq!(parsed, hdr); + assert_eq!(parsed.signature.bytes(), EWF1_EVF_SIGNATURE); + } + + #[test] + fn test_roundtrip_lvf_bytes() { + let hdr = Ewf1FileHeader::new(Ewf1Signature::Lvf, 1); + let bytes = hdr.to_bytes(); + let parsed = Ewf1FileHeader::parse(&bytes).unwrap(); + assert_eq!(parsed, hdr); + assert_eq!(parsed.signature.bytes(), EWF1_LVF_SIGNATURE); + } + + #[test] + fn test_parse_rejects_unknown_signature() { + let mut bytes = [0u8; EWF1_FILE_HEADER_SIZE]; + bytes[0..8].copy_from_slice(b"NOTEVF1!"); + let err = Ewf1FileHeader::parse(&bytes).unwrap_err(); + assert!(matches!(err, Error::Invalid(_))); + } +} diff --git a/crates/ewf/src/ewf1/header.rs b/crates/ewf/src/ewf1/header.rs new file mode 100644 index 0000000..9b22569 --- /dev/null +++ b/crates/ewf/src/ewf1/header.rs @@ -0,0 +1,292 @@ +//! EWF1 header parsing primitives (`header`, `header2`) and EWFX `xheader`. +//! +//! EWF1 stores human-oriented acquisition metadata in two places: +//! +//! - The **`header`** section: zlib-compressed ASCII text (EnCase-style) with CRLF line endings. +//! - The **`header2`** section: zlib-compressed UTF-16LE text (often EnCase 4–7), with categories. +//! +//! EWF-X (“EWFX”) adds: +//! +//! - The **`xheader`** section: zlib-compressed UTF-8 XML that contains the header values as XML +//! elements (notably including *both* `acquiry_software` and `acquiry_software_version`). +//! +//! This module intentionally focuses on **spec-aligned extraction** into structured +//! [`crate::metadata::HeaderValues`]. Presentation/formatting (labels, date formatting) belongs to +//! binaries such as `ewfinfo`. +//! +//! References: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - “Header section” +//! - “Header2 values” +//! - “EWF-X” → “Xheader” +//! - libewf implementation: +//! - `external/libewf/libewf/libewf_header_values.c` + +use crate::metadata::HeaderValues; +use crate::{Error, Result}; + +/// A tag identifier used in EWF1 header strings. +/// +/// This enum is used by both ASCII `header` and UTF-16LE `header2` parsing. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf1HeaderTag { + /// `a` — unique description / description. + Description, + /// `c` — case number. + CaseNumber, + /// `n` — evidence number. + EvidenceNumber, + /// `e` — examiner name. + ExaminerName, + /// `t` — notes. + Notes, + /// `av` — acquisition software **version** (e.g. EnCase version used to acquire media). + AcquisitionSoftwareVersion, + /// `ov` — acquisition platform/operating system. + AcquisitionOperatingSystem, + /// `m` — acquisition date/time string. + AcquisitionDateTime, + /// `u` — system date/time string. + SystemDateTime, + /// `p` — password hash (or `0` if not set). + PasswordHash, + /// `r` — compression level indicator (ASCII header only; not currently surfaced in metadata). + CompressionLevel, + /// `md` — media model (header2 only; not currently surfaced in metadata). + Model, + /// `sn` — serial number (header2 only; not currently surfaced in metadata). + SerialNumber, + /// Unknown/unsupported tag. + Unknown, +} + +impl Ewf1HeaderTag { + /// Parse an EWF1 header tag string into a typed identifier. + pub(crate) fn parse(tag: &str) -> Self { + match tag { + "a" => Self::Description, + "c" => Self::CaseNumber, + "n" => Self::EvidenceNumber, + "e" => Self::ExaminerName, + "t" => Self::Notes, + "av" => Self::AcquisitionSoftwareVersion, + "ov" => Self::AcquisitionOperatingSystem, + "m" => Self::AcquisitionDateTime, + "u" => Self::SystemDateTime, + "p" => Self::PasswordHash, + "r" => Self::CompressionLevel, + "md" => Self::Model, + "sn" => Self::SerialNumber, + _ => Self::Unknown, + } + } + + fn apply(self, value: String, out: &mut HeaderValues) { + match self { + Self::Description => out.description = Some(value), + Self::CaseNumber => out.case_number = Some(value), + Self::EvidenceNumber => out.evidence_number = Some(value), + Self::ExaminerName => out.examiner_name = Some(value), + Self::Notes => out.notes = Some(value), + Self::AcquisitionSoftwareVersion => out.acquisition_software_version = Some(value), + Self::AcquisitionOperatingSystem => out.acquisition_os = Some(value), + Self::AcquisitionDateTime => out.acquisition_datetime = Some(value), + Self::SystemDateTime => out.system_datetime = Some(value), + Self::PasswordHash => { + // libewf uses the literal character '0' to indicate “not set”. + if value != "0" { + out.password = Some(value); + } + } + Self::CompressionLevel | Self::Model | Self::SerialNumber | Self::Unknown => { + // Not surfaced in `HeaderValues` at the moment. + } + } + } +} + +/// Parser for an EWF1 `header` section (zlib-decompressed ASCII text). +#[derive(Debug, Clone, Copy)] +pub(crate) struct Ewf1HeaderAscii<'a> { + decompressed: &'a [u8], +} + +impl<'a> Ewf1HeaderAscii<'a> { + pub(crate) fn new(decompressed: &'a [u8]) -> Self { + Self { decompressed } + } + + pub(crate) fn parse_into(self, out: &mut HeaderValues) -> Result<()> { + let s = String::from_utf8_lossy(self.decompressed); + let mut lines = s.lines(); + + let _category_count = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header missing category count".to_string()))?; + let _category = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header missing category name".to_string()))?; + + let tags_line = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header missing tags line".to_string()))?; + let values_line = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header missing values line".to_string()))?; + + let tags: Vec<&str> = tags_line.trim_end_matches('\r').split('\t').collect(); + let values: Vec<&str> = values_line.trim_end_matches('\r').split('\t').collect(); + + if tags.len() != values.len() { + return Err(Error::Invalid( + "EWF1 header tags/values column count mismatch".to_string(), + )); + } + + for (t, v) in tags.into_iter().zip(values) { + Ewf1HeaderTag::parse(t).apply(v.to_string(), out); + } + + Ok(()) + } +} + +/// Parser for an EWF1 `header2` section (zlib-decompressed UTF-16LE text). +#[derive(Debug, Clone, Copy)] +pub(crate) struct Ewf1Header2Utf16Le<'a> { + decompressed: &'a [u8], +} + +impl<'a> Ewf1Header2Utf16Le<'a> { + pub(crate) fn new(decompressed: &'a [u8]) -> Self { + Self { decompressed } + } + + pub(crate) fn parse_into(self, out: &mut HeaderValues) -> Result<()> { + let mut bytes = self.decompressed; + if bytes.len() >= 2 && bytes[0..2] == [0xff, 0xfe] { + bytes = &bytes[2..]; + } + if !bytes.len().is_multiple_of(2) { + return Err(Error::Invalid( + "EWF1 header2 UTF-16LE has odd byte length".to_string(), + )); + } + let u16s: Vec = bytes + .chunks_exact(2) + .map(|c| u16::from_le_bytes(c.try_into().expect("len=2"))) + .collect(); + let s = String::from_utf16_lossy(&u16s); + + let mut lines = s.lines(); + let _category_count = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header2 missing category count".to_string()))?; + let _category = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header2 missing category name".to_string()))?; + let tags_line = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header2 missing tags line".to_string()))?; + let values_line = lines + .next() + .ok_or_else(|| Error::Invalid("EWF1 header2 missing values line".to_string()))?; + + let tags: Vec<&str> = tags_line.trim_end_matches('\r').split('\t').collect(); + let values: Vec<&str> = values_line.trim_end_matches('\r').split('\t').collect(); + if tags.len() != values.len() { + return Err(Error::Invalid( + "EWF1 header2 tags/values column count mismatch".to_string(), + )); + } + + for (t, v) in tags.into_iter().zip(values) { + Ewf1HeaderTag::parse(t).apply(v.to_string(), out); + } + + Ok(()) + } +} + +/// Parser for an EWFX `xheader` section (zlib-decompressed UTF-8 XML). +#[derive(Debug, Clone, Copy)] +pub(crate) struct EwfxXHeaderXml<'a> { + decompressed: &'a [u8], +} + +impl<'a> EwfxXHeaderXml<'a> { + pub(crate) fn new(decompressed: &'a [u8]) -> Self { + Self { decompressed } + } + + pub(crate) fn parse_into(self, out: &mut HeaderValues) -> Result<()> { + use quick_xml::Reader; + use quick_xml::events::Event; + use std::io::Cursor; + + let mut reader = Reader::from_reader(Cursor::new(self.decompressed)); + reader.config_mut().trim_text(true); + + let mut buf: Vec = Vec::new(); + let mut current_tag: Option> = None; + + loop { + match reader.read_event_into(&mut buf) { + Ok(Event::Start(e)) => { + current_tag = Some(e.name().as_ref().to_vec()); + } + Ok(Event::End(_)) => { + current_tag = None; + } + Ok(Event::Text(e)) => { + let Some(tag) = current_tag.as_deref() else { + buf.clear(); + continue; + }; + + let decoded = e + .decode() + .map_err(|e| Error::Invalid(format!("invalid xheader XML: {e}")))?; + let text = quick_xml::escape::unescape(decoded.as_ref()) + .map_err(|e| Error::Invalid(format!("invalid xheader XML: {e}")))? + .into_owned(); + + match tag { + b"case_number" => out.case_number = Some(text), + b"evidence_number" => out.evidence_number = Some(text), + b"description" => out.description = Some(text), + b"examiner_name" => out.examiner_name = Some(text), + b"notes" => out.notes = Some(text), + b"acquiry_date" => out.acquisition_datetime = Some(text), + b"system_date" => out.system_datetime = Some(text), + b"acquiry_operating_system" => out.acquisition_os = Some(text), + b"acquiry_software" => out.acquisition_software = Some(text), + b"acquiry_software_version" => { + out.acquisition_software_version = Some(text) + } + _ => {} + } + } + Ok(Event::Eof) => break, + Ok(_) => {} + Err(e) => return Err(Error::Invalid(format!("invalid xheader XML: {e}"))), + } + buf.clear(); + } + + Ok(()) + } +} + +/// Normalizes EWF1 header values for libewf-compatible `ewfinfo` output. +/// +/// Some tooling populates `acquiry_software` with the same value as +/// `acquiry_software_version`. libewf will only print a single “Software version used” line in +/// that case, so we drop the redundant “Software used” field. +pub(crate) fn normalize_header_values(values: &mut HeaderValues) { + if values.acquisition_software.is_some() + && values.acquisition_software == values.acquisition_software_version + { + values.acquisition_software = None; + } +} diff --git a/crates/ewf/src/ewf1/mod.rs b/crates/ewf/src/ewf1/mod.rs new file mode 100644 index 0000000..419e2a5 --- /dev/null +++ b/crates/ewf/src/ewf1/mod.rs @@ -0,0 +1,29 @@ +//! EWF1 (E01/S01/L01) format primitives. +//! +//! This module contains small, **spec-driven** building blocks shared by the reader and writer +//! implementations for the original EWF container formats (EWF1). +//! +//! Reference material: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - libewf reference implementation: `external/libewf/` + +/// EWF1 EVF segment file signature (`EVF\t\r\n\xff\x00`). +pub(crate) const EWF1_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; + +/// EWF1 LVF segment file signature (`LVF\t\r\n\xff\x00`) used by logical evidence (`.L01`). +pub(crate) const EWF1_LVF_SIGNATURE: [u8; 8] = [0x4c, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; + +/// EWF1 file header size (13 bytes). +pub(crate) const EWF1_FILE_HEADER_SIZE: usize = 8 + 1 + 2 + 2; + +/// EWF1 section descriptor size (76 bytes). +pub(crate) const EWF1_SECTION_DESCRIPTOR_SIZE: usize = 16 + 8 + 8 + 40 + 4; + +/// EWF1 chunk table header size (24 bytes). +pub(crate) const EWF1_TABLE_HEADER_SIZE: usize = 4 + 4 + 8 + 4 + 4; + +pub(crate) mod file_header; +pub(crate) mod header; +pub(crate) mod runs; +pub(crate) mod section; +pub(crate) mod volume; diff --git a/crates/ewf/src/ewf1/runs.rs b/crates/ewf/src/ewf1/runs.rs new file mode 100644 index 0000000..ba0a833 --- /dev/null +++ b/crates/ewf/src/ewf1/runs.rs @@ -0,0 +1,124 @@ +//! EWF1 run-list sections (`session`, `error2`). +//! +//! These sections encode lists of sector runs: +//! - `session`: optical media session boundaries (start sectors). +//! - `error2`: acquisition read error sector ranges. +//! +//! References: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - libewf implementation: +//! - `external/libewf/libewf/libewf_session_section.c` +//! - `external/libewf/libewf/libewf_error2_section.c` + +use crate::metadata::SectorRun; +use crate::{Error, Result}; +use binrw::{BinRead as _, binrw}; +use std::io::Cursor; + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1SessionSection { + count: u32, + #[brw(pad_after = 32)] + _header_padding: (), + #[br(count = count)] + entries: Vec, + checksum: u32, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1SessionEntry { + _unknown0: u32, + start_sector: u32, + #[brw(pad_after = 24)] + _reserved: (), +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1Error2Section { + count: u32, + #[brw(pad_after = 516)] + _header_padding: (), + #[br(count = count)] + entries: Vec, + checksum: u32, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1Error2Entry { + start_sector: u32, + sector_count: u32, +} + +/// EWF1 `session` section parser. +#[derive(Debug, Clone, Copy)] +pub(crate) struct Ewf1SessionSection<'a> { + data: &'a [u8], +} + +impl<'a> Ewf1SessionSection<'a> { + pub(crate) fn new(data: &'a [u8]) -> Self { + Self { data } + } + + /// Parse session runs. + /// + /// The section stores a list of session start sectors; the run length is derived from the next + /// start (or from `total_sectors` for the last run). + pub(crate) fn runs(self, total_sectors: u64) -> Result> { + let raw = RawEwf1SessionSection::read(&mut Cursor::new(self.data)) + .map_err(|e| Error::Invalid(format!("invalid EWF1 session section: {e}")))?; + + let mut starts: Vec = Vec::with_capacity(raw.entries.len()); + for e in raw.entries { + starts.push(u64::from(e.start_sector)); + } + + let mut runs: Vec = Vec::with_capacity(starts.len()); + for (i, start) in starts.iter().copied().enumerate() { + let sector_count = if let Some(next) = starts.get(i + 1).copied() { + next.saturating_sub(start) + } else { + total_sectors.saturating_sub(start) + }; + runs.push(SectorRun { + start_sector: start, + sector_count, + }); + } + Ok(runs) + } +} + +/// EWF1 `error2` section parser. +#[derive(Debug, Clone, Copy)] +pub(crate) struct Ewf1Error2Section<'a> { + data: &'a [u8], +} + +impl<'a> Ewf1Error2Section<'a> { + pub(crate) fn new(data: &'a [u8]) -> Self { + Self { data } + } + + pub(crate) fn runs(self) -> Result> { + let raw = RawEwf1Error2Section::read(&mut Cursor::new(self.data)) + .map_err(|e| Error::Invalid(format!("invalid EWF1 error2 section: {e}")))?; + + Ok(raw + .entries + .into_iter() + .map(|e| SectorRun { + start_sector: u64::from(e.start_sector), + sector_count: u64::from(e.sector_count), + }) + .collect()) + } +} diff --git a/crates/ewf/src/ewf1/section.rs b/crates/ewf/src/ewf1/section.rs new file mode 100644 index 0000000..38634e4 --- /dev/null +++ b/crates/ewf/src/ewf1/section.rs @@ -0,0 +1,302 @@ +//! EWF1 section descriptor parsing. +//! +//! EWF1 stores each section as: +//! +//! - a 76-byte **section descriptor** (type string, offsets, size, Adler-32), +//! - followed by the section **body** (raw bytes, possibly zlib-compressed depending on section), +//! - and then the next section descriptor begins (unless this is the last section). +//! +//! The section descriptor's `size` field is the total section size **including** the descriptor. +//! Some writers (notably for the `next`/`done` sections) set `size = 0` and rely on `next_offset` +//! instead; libewf infers the size from `next_offset` in that case. +//! +//! References: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - libewf implementation: +//! - `external/libewf/libewf/libewf_section_descriptor.c` + +use std::fs::File; +use std::io; + +use crate::util::{adler32_rfc1950, parse_ascii_nul_terminated, read_exact_at}; +use crate::{Error, Result}; +use binrw::{BinRead as _, BinWrite as _, binrw}; +use std::io::Cursor; + +use super::{EWF1_SECTION_DESCRIPTOR_SIZE, EWF1_TABLE_HEADER_SIZE}; + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy)] +struct RawEwf1SectionDescriptor { + type_bytes: [u8; 16], + next_offset: u64, + size: u64, + #[brw(pad_after = 40)] + _reserved: (), + checksum: u32, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy)] +struct RawEwf1TableHeader { + number_of_entries: u32, + #[brw(pad_after = 4)] + _reserved0: (), + base_offset: u64, + #[brw(pad_after = 4)] + _reserved1: (), + checksum: u32, +} + +/// Known EWF1 section type strings. +#[derive(Debug, Clone, PartialEq, Eq)] +pub(crate) enum Ewf1SectionType { + Header, + Header2, + Volume, + Disk, + Data, + Sectors, + Sector, + Table, + Table2, + Digest, + Hash, + Session, + Error2, + LTree, + XHeader, + XHash, + Next, + Done, + Unknown(String), +} + +impl Ewf1SectionType { + pub(crate) fn parse(value: &str) -> Self { + match value { + "header" => Self::Header, + "header2" => Self::Header2, + "volume" => Self::Volume, + "disk" => Self::Disk, + "data" => Self::Data, + "sectors" => Self::Sectors, + "sector" => Self::Sector, + "table" => Self::Table, + "table2" => Self::Table2, + "digest" => Self::Digest, + "hash" => Self::Hash, + "session" => Self::Session, + "error2" => Self::Error2, + "ltree" => Self::LTree, + "xheader" => Self::XHeader, + "xhash" => Self::XHash, + "next" => Self::Next, + "done" => Self::Done, + other => Self::Unknown(other.to_string()), + } + } + + pub(crate) fn as_str(&self) -> &str { + match self { + Self::Header => "header", + Self::Header2 => "header2", + Self::Volume => "volume", + Self::Disk => "disk", + Self::Data => "data", + Self::Sectors => "sectors", + Self::Sector => "sector", + Self::Table => "table", + Self::Table2 => "table2", + Self::Digest => "digest", + Self::Hash => "hash", + Self::Session => "session", + Self::Error2 => "error2", + Self::LTree => "ltree", + Self::XHeader => "xheader", + Self::XHash => "xhash", + Self::Next => "next", + Self::Done => "done", + Self::Unknown(s) => s, + } + } + + pub(crate) fn is_terminal(&self) -> bool { + matches!(self, Self::Next | Self::Done) + } +} + +/// An EWF1 section descriptor. +#[derive(Debug, Clone)] +pub(crate) struct Ewf1SectionDescriptor { + /// Offset of the section descriptor (relative to the start of the segment file). + pub(crate) start_offset: u64, + /// Parsed section type. + pub(crate) section_type: Ewf1SectionType, + /// Total section size in bytes, including the descriptor. + pub(crate) size: u64, +} + +impl Ewf1SectionDescriptor { + /// Parse a section descriptor at an absolute file offset. + pub(crate) fn parse_at(file: &File, file_len: u64, start_offset: u64) -> Result { + if start_offset >= file_len { + return Err(io::Error::from(io::ErrorKind::UnexpectedEof).into()); + } + + let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; + read_exact_at(file, start_offset, &mut raw)?; + + let stored = u32::from_le_bytes( + raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..] + .try_into() + .expect("len=4"), + ); + let calculated = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); + if stored != calculated { + return Err(Error::Corrupt( + "section descriptor checksum mismatch".to_string(), + )); + } + + let raw_desc = RawEwf1SectionDescriptor::read(&mut Cursor::new(&raw[..])).map_err(|e| { + Error::Invalid(format!("invalid EWF1 section descriptor encoding: {e}")) + })?; + + let type_string = parse_ascii_nul_terminated(&raw_desc.type_bytes); + let section_type = Ewf1SectionType::parse(&type_string); + + let next_offset = raw_desc.next_offset; + let mut size = raw_desc.size; + + // libewf behavior: some writers leave `size = 0`, but set `next_offset`; infer size from that. + if size == 0 && next_offset != start_offset && next_offset >= start_offset { + size = next_offset - start_offset; + } + + Ok(Self { + start_offset, + section_type, + size, + }) + } + + /// Returns the byte range containing the section body (`[start, end)`). + pub(crate) fn data_range(&self) -> Result<(u64, u64)> { + let start = self + .start_offset + .checked_add(EWF1_SECTION_DESCRIPTOR_SIZE as u64) + .ok_or_else(|| Error::Invalid("section range overflow".to_string()))?; + let end = self + .start_offset + .checked_add(self.size) + .ok_or_else(|| Error::Invalid("section range overflow".to_string()))?; + Ok((start, end)) + } + + /// Returns how far the scanning cursor should advance after this descriptor. + pub(crate) fn advance_len(&self) -> Result { + let advance = if self.size != 0 { + self.size + } else { + // libewf: for last sections (`next`/`done`) some writers set size=0; advance by descriptor size. + EWF1_SECTION_DESCRIPTOR_SIZE as u64 + }; + + if advance == 0 { + return Err(Error::Invalid( + "zero advance while scanning sections".to_string(), + )); + } + Ok(advance) + } + + /// Scan section descriptors from `first_section_offset` until the terminal section. + pub(crate) fn scan(file: &File, file_len: u64, first_section_offset: u64) -> Result> { + let mut sections = Vec::new(); + let mut offset = first_section_offset; + + // Hard safety cap: avoid pathological scans on corrupted inputs. + for _ in 0..100_000 { + if offset == 0 || offset >= file_len { + break; + } + + let desc = Self::parse_at(file, file_len, offset)?; + let is_last = desc.section_type.is_terminal(); + let advance = desc.advance_len()?; + + sections.push(desc); + if is_last { + break; + } + + offset = offset.saturating_add(advance); + } + + if sections.is_empty() { + return Err(Error::Invalid("no EWF sections found".to_string())); + } + + Ok(sections) + } +} + +/// Writes a canonical EWF1 section descriptor to bytes. +/// +/// This is shared by the writer (to construct descriptors) and tests/tools. The `start_offset` +/// argument is accepted for API parity with libewf-like helpers, but is not encoded in the +/// descriptor itself. +pub(crate) fn make_ewf1_section_descriptor( + type_string: &str, + _start_offset: u64, + next_offset: u64, + size: u64, +) -> [u8; EWF1_SECTION_DESCRIPTOR_SIZE] { + let mut type_bytes = [0u8; 16]; + let src = type_string.as_bytes(); + let copy_len = src.len().min(type_bytes.len().saturating_sub(1)); + type_bytes[..copy_len].copy_from_slice(&src[..copy_len]); + + let desc = RawEwf1SectionDescriptor { + type_bytes, + next_offset, + size, + _reserved: (), + checksum: 0, + }; + + let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; + desc.write(&mut Cursor::new(&mut raw[..])) + .expect("in-memory write cannot fail"); + + let checksum = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); + raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); + raw +} + +/// Writes a canonical EWF1 chunk table header to bytes. +/// +/// This is used by both EWF-E01 (`table`/`table2`) and EWF-S01 (`table`) table sections. +pub(crate) fn make_ewf1_table_header( + number_of_entries: u32, + base_offset: u64, +) -> [u8; EWF1_TABLE_HEADER_SIZE] { + let hdr = RawEwf1TableHeader { + number_of_entries, + _reserved0: (), + base_offset, + _reserved1: (), + checksum: 0, + }; + + let mut out = [0u8; EWF1_TABLE_HEADER_SIZE]; + hdr.write(&mut Cursor::new(&mut out[..])) + .expect("in-memory write cannot fail"); + + let checksum = adler32_rfc1950(&out[..EWF1_TABLE_HEADER_SIZE - 4]); + out[EWF1_TABLE_HEADER_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); + out +} diff --git a/crates/ewf/src/ewf1/volume.rs b/crates/ewf/src/ewf1/volume.rs new file mode 100644 index 0000000..af719c8 --- /dev/null +++ b/crates/ewf/src/ewf1/volume.rs @@ -0,0 +1,466 @@ +//! EWF1 `volume` / `disk` / `data` section parsing (metadata). +//! +//! This module is intentionally small and **format-focused**: it parses the raw bytes of the +//! “volume-like” section body into typed enums and values, without embedding display strings in the +//! parser itself. +//! +//! Reference material: +//! - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` +//! - “Volume section” (94-byte “EWF specification” variant; 1052-byte “EnCase/FTK” variant) +//! +//! ## Input format +//! +//! Callers must pass the **raw section body** bytes as stored in the EWF1 container (i.e. the +//! `data_range()` portion of an EWF1 section descriptor). +//! +//! The layout is a “volume-like” header used by multiple section types (`volume`, `disk`, and +//! sometimes `data`), with additional fields present in the EnCase/FTK “1052-byte volume” variant: +//! +//! - Offset `0x00` (`u8`): media type code +//! - Offset `0x08` (`u32` LE): sectors per chunk +//! - Offset `0x0c` (`u32` LE): bytes per sector +//! - Offset `0x10` (`u32` or `u64` LE): number of sectors +//! - If the buffer is at least 24 bytes: interpret as `u64` at `0x10..0x18` +//! - Otherwise: interpret as `u32` at `0x10..0x14` +//! - Offset `0x24` (`u8`): media flags (1052-byte variant; optional) +//! - Offset `0x34` (`u8`): compression level hint (1052-byte variant; optional) +//! - Offset `0x40..0x50` (`[u8; 16]`): set identifier (1052-byte variant; optional) +//! +//! Notes: +//! - This parser only models fields that are required for the `ewfinfo` report. +//! - Unknown codes are preserved via `Unknown(u8)` enum variants. + +use std::fmt; + +use crate::util::adler32_rfc1950; +use crate::{Error, Result}; +use binrw::{BinRead as _, BinWrite as _, binrw}; +use bitflags::bitflags; +use std::io::{Cursor, SeekFrom}; + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1VolumeLikeSectionBody { + /// Byte 0: media type code (in the 1052-byte variant). + /// + /// In the 94-byte “spec” variant, this is part of a 4-byte reserved field that commonly holds + /// `0x01 0x00 0x00 0x00`, which coincides with the “fixed disk” media type code. + media_type_code: u8, + #[brw(pad_after = 3)] + _reserved0: (), + + _chunk_count: u32, + sectors_per_chunk: u32, + bytes_per_sector: u32, + + // EWF1 stores the sector count as either 32-bit (older variants) or 64-bit (newer ones). + number_of_sectors_low: u32, + #[br(try)] + number_of_sectors_high: Option, + + // Fields below are specific to the 1052-byte EnCase/FTK/linen variant. We read them with + // explicit offsets and tolerate missing data so this parser can operate on truncated inputs. + #[br(try, seek_before = SeekFrom::Start(36))] + media_flags_raw: Option, + #[br(try, seek_before = SeekFrom::Start(52))] + compression_level_raw: Option, + #[br(try, seek_before = SeekFrom::Start(56))] + error_granularity: Option, + #[br(try, seek_before = SeekFrom::Start(64))] + set_identifier_raw: Option<[u8; 16]>, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1VolumeSectionE01_1052 { + media_type_code: u8, + #[brw(pad_after = 3)] + _reserved0: (), + chunk_count: u32, + sectors_per_chunk: u32, + bytes_per_sector: u32, + number_of_sectors: u64, + + #[brw(pad_after = 12)] + _reserved1: (), + media_flags: u8, + + #[brw(pad_after = 15)] + _reserved2: (), + compression_level: u8, + + #[brw(pad_after = 3)] + _reserved3: (), + error_granularity: u32, + + #[brw(pad_after = 4)] + _reserved4: (), + #[brw(pad_after = 968)] + set_identifier: [u8; 16], + + checksum: u32, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone)] +struct RawEwf1VolumeSectionS01_94 { + reserved0_low: u8, + #[brw(pad_after = 3)] + _reserved0: (), + chunk_count: u32, + sectors_per_chunk: u32, + bytes_per_sector: u32, + number_of_sectors: u32, + + #[brw(pad_after = 65)] + _reserved1: (), + #[brw(magic = b"SMART")] + _smart_signature: (), + checksum: u32, +} + +/// EWF1 media type code (byte 0 of the volume-like section body). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf1MediaType { + RemovableDisk, + FixedDisk, + OpticalDisk, + SingleFiles, + MemoryRam, + Unknown(u8), +} + +impl Ewf1MediaType { + pub(crate) fn from_code(code: u8) -> Self { + match code { + 0x00 => Self::RemovableDisk, + 0x01 => Self::FixedDisk, + 0x03 => Self::OpticalDisk, + 0x0e => Self::SingleFiles, + 0x10 => Self::MemoryRam, + other => Self::Unknown(other), + } + } +} + +impl fmt::Display for Ewf1MediaType { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + let s = match self { + Ewf1MediaType::RemovableDisk => "removable disk", + Ewf1MediaType::FixedDisk => "fixed disk", + Ewf1MediaType::OpticalDisk => "optical disk (CD/DVD/BD)", + Ewf1MediaType::SingleFiles => "single files", + Ewf1MediaType::MemoryRam => "memory (RAM)", + Ewf1MediaType::Unknown(_) => "unknown", + }; + f.write_str(s) + } +} + +bitflags! { + /// EWF1 media flags byte (offset 36 in the 1052-byte variant). + /// + /// Unknown bits are preserved (see [`Ewf1MediaFlags::from_raw`]). + #[derive(Debug, Clone, Copy, PartialEq, Eq)] + pub(crate) struct Ewf1MediaFlags: u8 { + /// libewf: `LIBEWF_MEDIA_FLAG_PHYSICAL` + const PHYSICAL = 0x02; + } +} + +impl Ewf1MediaFlags { + pub(crate) fn from_raw(raw: u8) -> Self { + Self::from_bits_retain(raw) + } + + pub(crate) fn is_physical(self) -> bool { + self.contains(Self::PHYSICAL) + } +} + +/// Compression level hint byte (offset 52 in the 1052-byte variant). +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf1VolumeCompressionLevel { + NoCompression, + GoodFastCompression, + BestCompression, + Unknown(u8), + NotRecorded, +} + +impl Ewf1VolumeCompressionLevel { + pub(crate) fn from_optional_code(code: Option) -> Self { + match code { + Some(0x00) => Self::NoCompression, + Some(0x01) => Self::GoodFastCompression, + Some(0x02) => Self::BestCompression, + Some(other) => Self::Unknown(other), + None => Self::NotRecorded, + } + } +} + +impl fmt::Display for Ewf1VolumeCompressionLevel { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + let s = match self { + Ewf1VolumeCompressionLevel::NoCompression => "no compression", + Ewf1VolumeCompressionLevel::GoodFastCompression => "good (fast) compression", + Ewf1VolumeCompressionLevel::BestCompression => "best compression", + Ewf1VolumeCompressionLevel::Unknown(_) | Ewf1VolumeCompressionLevel::NotRecorded => { + "unknown compression" + } + }; + f.write_str(s) + } +} + +/// Parsed EWF1 volume-like section fields used by `ewfinfo`. +#[derive(Debug, Clone, PartialEq, Eq)] +pub(crate) struct Ewf1VolumeInfo { + pub(crate) sectors_per_chunk: u32, + pub(crate) error_granularity: u32, + pub(crate) bytes_per_sector: u32, + pub(crate) number_of_sectors: u64, + pub(crate) media_size: u64, + pub(crate) compression_level: Ewf1VolumeCompressionLevel, + pub(crate) set_identifier: Option<[u8; 16]>, + pub(crate) media_type: Ewf1MediaType, + pub(crate) is_physical: bool, +} + +impl Ewf1VolumeInfo { + /// Parse an EWF1 “volume-like” section **body**. + /// + /// The input must be the raw bytes of the `volume`/`disk`/`data` section body (not including + /// the 76-byte section descriptor). + pub(crate) fn parse_from_volume_like_section_body(data: &[u8]) -> Result { + if data.len() < 20 { + return Err(Error::Invalid( + "short EWF1 volume-like section body".to_string(), + )); + } + + let raw = RawEwf1VolumeLikeSectionBody::read(&mut Cursor::new(data)) + .map_err(|e| Error::Invalid(format!("invalid EWF1 volume-like section body: {e}")))?; + + // The 1052-byte EnCase/FTK/linen variant stores `media_type` at byte 0. + // The 94-byte “EWF specification” variant uses a 4-byte reserved field at offset 0 that + // *contains* 0x01, which coincides with the “fixed disk” media type code. We keep that + // behavior for compatibility. + let media_type = Ewf1MediaType::from_code(raw.media_type_code); + + // NOTE: The `chunk_count` field exists at 0x04, but is currently unused by `ewfinfo`. + // We still parse the geometry fields and sector count to compute media_size. + let sectors_per_chunk = raw.sectors_per_chunk; + let bytes_per_sector = raw.bytes_per_sector; + + if sectors_per_chunk == 0 || bytes_per_sector == 0 { + return Err(Error::Invalid("invalid EWF1 volume parameters".to_string())); + } + + // In EWF1 the sector count is 32-bit in older variants and 64-bit in newer ones. + let number_of_sectors = if let Some(high) = raw.number_of_sectors_high { + (u64::from(high) << 32) | u64::from(raw.number_of_sectors_low) + } else { + u64::from(raw.number_of_sectors_low) + }; + + let media_size = number_of_sectors + .checked_mul(bytes_per_sector as u64) + .ok_or_else(|| Error::Invalid("media size overflow".to_string()))?; + + let is_1052_variant = data.len() >= 1052; + + // Media flags live at offset 36 in the 1052-byte EnCase/FTK/linen volume variant. + let media_flags = if is_1052_variant { + raw.media_flags_raw.map(Ewf1MediaFlags::from_raw) + } else { + None + }; + let is_physical = media_flags.map(|f| f.is_physical()).unwrap_or(false); + + // Sector error granularity lives at offset 56 in the 1052-byte EnCase/FTK/linen volume variant. + let error_granularity = if is_1052_variant { + raw.error_granularity.unwrap_or(0) + } else { + 0 + }; + + // Compression level at offset 52 for the 1052-byte variant. + let compression_level = if is_1052_variant { + Ewf1VolumeCompressionLevel::from_optional_code(raw.compression_level_raw) + } else { + Ewf1VolumeCompressionLevel::NotRecorded + }; + + // Set identifier is stored at [64..80] in the 1052-byte volume variant. + let set_identifier = if is_1052_variant { + raw.set_identifier_raw + .and_then(|id| (id.iter().any(|&b| b != 0)).then_some(id)) + } else { + None + }; + + Ok(Self { + sectors_per_chunk, + error_granularity, + bytes_per_sector, + number_of_sectors, + media_size, + compression_level, + set_identifier, + media_type, + is_physical, + }) + } +} + +pub(crate) fn build_volume_section_e01_1052( + chunk_count: u64, + sectors_per_chunk: u32, + error_granularity: u32, + bytes_per_sector: u32, + number_of_sectors: u64, + compression_level: u8, + set_identifier: [u8; 16], +) -> Vec { + // FTK Imager / EnCase 1–7 / linen volume (1052 bytes) variant. + // + // See `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc`: + // “Volume section” → “FTK Imager, EnCase 1 to 7 and linen 5 to 7 (EWF-E01)”. + let v = RawEwf1VolumeSectionE01_1052 { + media_type_code: 0x01, // fixed media + _reserved0: (), + chunk_count: chunk_count as u32, + sectors_per_chunk, + bytes_per_sector, + number_of_sectors, + _reserved1: (), + media_flags: 0x01, // “is an image file” + _reserved2: (), + compression_level, + _reserved3: (), + error_granularity, + _reserved4: (), + set_identifier, + checksum: 0, + }; + + let mut out = vec![0u8; 1052]; + v.write(&mut Cursor::new(&mut out[..])) + .expect("in-memory write cannot fail"); + let checksum = adler32_rfc1950(&out[..1048]).to_le_bytes(); + out[1048..1052].copy_from_slice(&checksum); + out +} + +pub(crate) fn build_volume_section_s01_94( + chunk_count: u64, + sectors_per_chunk: u32, + bytes_per_sector: u32, + number_of_sectors: u64, +) -> Vec { + // EWF specification (94 bytes) variant used by SMART (EWF-S01). + // + // See `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc`: + // “Volume section” → “EWF specification” and “SMART (EWF-S01)”. + let v = RawEwf1VolumeSectionS01_94 { + reserved0_low: 0x01, + _reserved0: (), + chunk_count: chunk_count as u32, + sectors_per_chunk, + bytes_per_sector, + number_of_sectors: number_of_sectors as u32, + _reserved1: (), + _smart_signature: (), + checksum: 0, + }; + + let mut out = vec![0u8; 94]; + v.write(&mut Cursor::new(&mut out[..])) + .expect("in-memory write cannot fail"); + let checksum = adler32_rfc1950(&out[..90]).to_le_bytes(); + out[90..94].copy_from_slice(&checksum); + out +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_parse_volume_like_section_body_64bit_sector_count() { + let mut data = [0u8; 24]; + data[0] = 0x01; // fixed disk + data[8..12].copy_from_slice(&64u32.to_le_bytes()); + data[12..16].copy_from_slice(&512u32.to_le_bytes()); + data[16..24].copy_from_slice(&2880u64.to_le_bytes()); + + let v = Ewf1VolumeInfo::parse_from_volume_like_section_body(&data).unwrap(); + assert_eq!(v.sectors_per_chunk, 64); + assert_eq!(v.error_granularity, 0); + assert_eq!(v.bytes_per_sector, 512); + assert_eq!(v.number_of_sectors, 2880); + assert_eq!(v.media_size, 2880u64 * 512); + assert_eq!(v.media_type.to_string(), "fixed disk"); + assert!(!v.is_physical); + assert_eq!(v.compression_level.to_string(), "unknown compression"); + assert!(v.set_identifier.is_none()); + } + + #[test] + fn test_parse_volume_like_section_body_32bit_sector_count() { + let mut data = [0u8; 20]; + data[0] = 0x00; // removable disk + data[8..12].copy_from_slice(&1u32.to_le_bytes()); + data[12..16].copy_from_slice(&512u32.to_le_bytes()); + data[16..20].copy_from_slice(&2880u32.to_le_bytes()); + + let v = Ewf1VolumeInfo::parse_from_volume_like_section_body(&data).unwrap(); + assert_eq!(v.number_of_sectors, 2880); + assert_eq!(v.media_type.to_string(), "removable disk"); + } + + #[test] + fn test_parse_volume_like_section_body_variant_fields_flags_compression_set_id() { + let mut data = vec![0u8; 1052]; + data[0] = 0x03; // optical disk + data[8..12].copy_from_slice(&32u32.to_le_bytes()); + data[12..16].copy_from_slice(&2048u32.to_le_bytes()); + data[16..24].copy_from_slice(&10u64.to_le_bytes()); + data[36] = 0x02; // physical flag + data[52] = 0x02; // best compression + data[56..60].copy_from_slice(&7u32.to_le_bytes()); // error granularity + + let set_id: [u8; 16] = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16]; + data[64..80].copy_from_slice(&set_id); + + let v = Ewf1VolumeInfo::parse_from_volume_like_section_body(&data).unwrap(); + assert_eq!(v.media_type.to_string(), "optical disk (CD/DVD/BD)"); + assert!(v.is_physical); + assert_eq!(v.compression_level.to_string(), "best compression"); + assert_eq!(v.error_granularity, 7); + assert_eq!(v.set_identifier, Some(set_id)); + } + + #[test] + fn test_parse_volume_like_section_body_rejects_zero_geometry() { + let mut data = [0u8; 24]; + data[8..12].copy_from_slice(&0u32.to_le_bytes()); + data[12..16].copy_from_slice(&512u32.to_le_bytes()); + data[16..24].copy_from_slice(&1u64.to_le_bytes()); + + let err = Ewf1VolumeInfo::parse_from_volume_like_section_body(&data).unwrap_err(); + assert!(matches!(err, Error::Invalid(_))); + } + + #[test] + fn test_parse_volume_like_section_body_rejects_too_short() { + let data = [0u8; 19]; + let err = Ewf1VolumeInfo::parse_from_volume_like_section_body(&data).unwrap_err(); + assert!(matches!(err, Error::Invalid(_))); + } +} diff --git a/crates/ewf/src/ewf2/chunk.rs b/crates/ewf/src/ewf2/chunk.rs new file mode 100644 index 0000000..83f6c13 --- /dev/null +++ b/crates/ewf/src/ewf2/chunk.rs @@ -0,0 +1,43 @@ +//! EWF2 chunk-table primitives. +//! +//! EWF2 stores per-chunk metadata in the *sector table* section. Each table entry contains a flags +//! bitmask describing how the corresponding chunk is stored (compressed, checksumed, pattern fill, +//! etc.). + +use bitflags::bitflags; + +bitflags! { + /// EWF2 sector-table entry flags. + /// + /// Unknown bits are preserved (see [`Ewf2ChunkDataFlags::from_raw`]). + #[derive(Debug, Clone, Copy, PartialEq, Eq)] + pub(crate) struct Ewf2ChunkDataFlags: u32 { + const COMPRESSED = 0x0000_0001; + const CHECKSUMED = 0x0000_0002; + const PATTERNFILL = 0x0000_0004; + } +} + +impl Ewf2ChunkDataFlags { + pub(crate) fn from_raw(raw: u32) -> Self { + Self::from_bits_retain(raw) + } + + pub(crate) fn raw(self) -> u32 { + self.bits() + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_chunk_data_flags_preserves_unknown_bits() { + let raw = 0x4000_0000u32 | Ewf2ChunkDataFlags::CHECKSUMED.bits(); + let flags = Ewf2ChunkDataFlags::from_raw(raw); + assert_eq!(flags.raw(), raw); + assert!(flags.contains(Ewf2ChunkDataFlags::CHECKSUMED)); + assert!(!flags.contains(Ewf2ChunkDataFlags::COMPRESSED)); + } +} diff --git a/crates/ewf/src/ewf2/file_header.rs b/crates/ewf/src/ewf2/file_header.rs new file mode 100644 index 0000000..c81c4a3 --- /dev/null +++ b/crates/ewf/src/ewf2/file_header.rs @@ -0,0 +1,128 @@ +//! EWF2 segment file header (Ex01/Lx01). +//! +//! Every EWF2 segment starts with a fixed-size 32-byte header. +//! +//! ## Layout (EWF2 2.1) +//! +//! The file header is 32 bytes of size and consists of: +//! +//! - Offset `0x00..0x08` (`[u8; 8]`): signature (`EVF2\r\n\x81\x00` or `LEF2\r\n\x81\x00`) +//! - Offset `0x08` (`u8`): major version (expected `2`) +//! - Offset `0x09` (`u8`): minor version (commonly `1`) +//! - Offset `0x0a..0x0c` (`u16` LE): compression method (see “Compression methods” in spec) +//! - Offset `0x0c..0x10` (`u32` LE): segment file number (series) +//! - Offset `0x10..0x20` (`[u8; 16]`): segment file set identifier (little-endian GUID v4) +//! +//! Reference material: +//! - `external/libewf/documentation/Expert Witness Compression Format 2 (EWF2).asciidoc` +//! - “Segment files” → “File header” + +use crate::{Error, Result}; +use binrw::{BinRead as _, BinWrite as _, binrw}; +use std::io::Cursor; + +pub(crate) const EWF2_FILE_HEADER_SIZE: usize = 32; + +/// EWF2-Ex01 signature (`EVF2\r\n\x81\x00`). +pub(crate) const EWF2_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x32, 0x0d, 0x0a, 0x81, 0x00]; + +/// EWF2-Lx01 signature (`LEF2\r\n\x81\x00`). +pub(crate) const EWF2_LEF_SIGNATURE: [u8; 8] = [0x4c, 0x45, 0x46, 0x32, 0x0d, 0x0a, 0x81, 0x00]; + +const EWF2_VERSION_MAJOR: u8 = 2; +const EWF2_VERSION_MINOR: u8 = 1; + +/// EWF2 container kind (Ex01 vs Lx01). +/// +/// The kind is encoded in the file header signature. +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf2Kind { + /// EWF2-Ex01 (`EVF2\r\n\x81\x00`) + #[brw(magic = b"EVF2\r\n\x81\0")] + Ex01, + /// EWF2-Lx01 (`LEF2\r\n\x81\x00`) + #[brw(magic = b"LEF2\r\n\x81\0")] + Lx01, +} + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) struct Ewf2FileHeader { + pub(crate) kind: Ewf2Kind, + #[br(assert( + major == EWF2_VERSION_MAJOR, + "unsupported EWF2 major version: {}", + major + ))] + pub(crate) major: u8, + pub(crate) minor: u8, + pub(crate) compression_method: u16, + pub(crate) segment_number: u32, + pub(crate) set_id: [u8; 16], +} + +impl Ewf2FileHeader { + pub(crate) fn new( + kind: Ewf2Kind, + compression_method: u16, + segment_number: u32, + set_id: [u8; 16], + ) -> Self { + Self { + kind, + major: EWF2_VERSION_MAJOR, + minor: EWF2_VERSION_MINOR, + compression_method, + segment_number, + set_id, + } + } + + pub(crate) fn parse(bytes: &[u8; EWF2_FILE_HEADER_SIZE]) -> Result { + let mut cur = Cursor::new(bytes.as_slice()); + let hdr = Self::read(&mut cur).map_err(|e| { + // Most failures here are signature/version mismatches, which callers expect as “invalid”. + Error::Invalid(format!("invalid EWF2 segment file header: {e}")) + })?; + Ok(hdr) + } + + pub(crate) fn to_bytes(self) -> [u8; EWF2_FILE_HEADER_SIZE] { + let mut out = [0u8; EWF2_FILE_HEADER_SIZE]; + self.write(&mut Cursor::new(&mut out[..])) + .expect("in-memory write cannot fail"); + out + } +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn test_roundtrip_ex01_bytes() { + let hdr = Ewf2FileHeader::new(Ewf2Kind::Ex01, 1, 7, [0x11; 16]); + let bytes = hdr.to_bytes(); + let parsed = Ewf2FileHeader::parse(&bytes).unwrap(); + assert_eq!(parsed, hdr); + } + + #[test] + fn test_roundtrip_lx01_bytes() { + let hdr = Ewf2FileHeader::new(Ewf2Kind::Lx01, 0, 1, [0x22; 16]); + let bytes = hdr.to_bytes(); + let parsed = Ewf2FileHeader::parse(&bytes).unwrap(); + assert_eq!(parsed, hdr); + } + + #[test] + fn test_parse_rejects_unknown_signature() { + let mut bytes = [0u8; EWF2_FILE_HEADER_SIZE]; + bytes[0..8].copy_from_slice(b"NOTEVF2!"); + let err = Ewf2FileHeader::parse(&bytes).unwrap_err(); + assert!(matches!(err, Error::Invalid(_))); + } +} diff --git a/crates/ewf/src/ewf2/mod.rs b/crates/ewf/src/ewf2/mod.rs new file mode 100644 index 0000000..942998a --- /dev/null +++ b/crates/ewf/src/ewf2/mod.rs @@ -0,0 +1,11 @@ +//! EWF2 (Ex01/Lx01) format primitives. +//! +//! This module contains small, **spec-driven** building blocks shared by the reader and writer +//! implementations. +//! +//! Reference material: +//! - `external/libewf/documentation/Expert Witness Compression Format 2 (EWF2).asciidoc` + +pub(crate) mod chunk; +pub(crate) mod file_header; +pub(crate) mod section; diff --git a/crates/ewf/src/ewf2/section.rs b/crates/ewf/src/ewf2/section.rs new file mode 100644 index 0000000..17ab727 --- /dev/null +++ b/crates/ewf/src/ewf2/section.rs @@ -0,0 +1,363 @@ +//! EWF2 section descriptor parsing. +//! +//! EWF2 stores each section as: +//! +//! - the section **data** (possibly compressed / encrypted), +//! - followed by a fixed-size **section descriptor** (64 bytes) *at the end of the section*. +//! +//! Section descriptors are chained backwards through the `previous_offset` field, so readers can +//! find all sections by starting at the last descriptor (at end-of-file) and walking backwards. +//! +//! References: +//! - `external/libewf/documentation/Expert Witness Compression Format 2 (EWF2).asciidoc` +//! - libewf implementation: +//! - `external/libewf/libewf/libewf_section_descriptor.c` + +use std::fs::File; +use std::io::Cursor; + +use crate::ewf2::file_header::EWF2_FILE_HEADER_SIZE; +use crate::util::{adler32_rfc1950, read_exact_at}; +use crate::{Error, Result}; +use binrw::{BinRead as _, BinWrite as _, binrw}; +use bitflags::bitflags; + +/// EWF2 section descriptor size (64 bytes). +pub(crate) const EWF2_SECTION_DESCRIPTOR_SIZE: usize = 64; + +#[binrw] +#[brw(little)] +#[derive(Debug, Clone, Copy)] +struct RawEwf2SectionDescriptor { + section_type: u32, + data_flags: u32, + previous_offset: u64, + data_size: u64, + descriptor_size: u32, + padding_size: u32, + data_integrity_hash: [u8; 16], + reserved: [u8; 12], + checksum: u32, +} + +/// EWF2 section type identifiers. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub(crate) enum Ewf2SectionType { + DeviceInformation, + CaseData, + SectorData, + SectorTable, + Md5Hash, + Sha1Hash, + Next, + Done, + SingleFilesData, + Unknown(u32), +} + +impl Ewf2SectionType { + pub(crate) fn from_u32(v: u32) -> Self { + match v { + 0x0000_0001 => Self::DeviceInformation, + 0x0000_0002 => Self::CaseData, + 0x0000_0003 => Self::SectorData, + 0x0000_0004 => Self::SectorTable, + 0x0000_0008 => Self::Md5Hash, + 0x0000_0009 => Self::Sha1Hash, + 0x0000_000d => Self::Next, + 0x0000_000f => Self::Done, + 0x0000_0020 => Self::SingleFilesData, + other => Self::Unknown(other), + } + } + + pub(crate) fn as_u32(self) -> u32 { + match self { + Self::DeviceInformation => 0x0000_0001, + Self::CaseData => 0x0000_0002, + Self::SectorData => 0x0000_0003, + Self::SectorTable => 0x0000_0004, + Self::Md5Hash => 0x0000_0008, + Self::Sha1Hash => 0x0000_0009, + Self::Next => 0x0000_000d, + Self::Done => 0x0000_000f, + Self::SingleFilesData => 0x0000_0020, + Self::Unknown(v) => v, + } + } +} + +bitflags! { + /// EWF2 section data flags. + /// + /// Unknown bits are preserved (see [`Ewf2SectionDataFlags::new`]). + #[derive(Debug, Clone, Copy, PartialEq, Eq)] + pub(crate) struct Ewf2SectionDataFlags: u32 { + const MD5_HASHED = 0x0000_0001; + const ENCRYPTED = 0x0000_0002; + } +} + +impl Ewf2SectionDataFlags { + /// Constructs flags from the raw on-disk bitmask, preserving unknown bits. + pub(crate) fn new(raw: u32) -> Self { + Self::from_bits_retain(raw) + } + + pub(crate) fn raw(self) -> u32 { + self.bits() + } + + pub(crate) fn has_md5_integrity_hash(self) -> bool { + self.contains(Self::MD5_HASHED) + } + + pub(crate) fn is_encrypted(self) -> bool { + self.contains(Self::ENCRYPTED) + } +} + +/// An EWF2 section descriptor plus derived section range information. +#[derive(Debug, Clone)] +pub(crate) struct Ewf2Section { + pub(crate) section_type: Ewf2SectionType, + pub(crate) data_flags: Ewf2SectionDataFlags, + pub(crate) previous_offset: u64, + + #[allow(dead_code)] + pub(crate) data_size: u64, + #[allow(dead_code)] + pub(crate) descriptor_size: u32, + pub(crate) padding_size: u32, + #[allow(dead_code)] + pub(crate) data_integrity_hash: [u8; 16], + + /// Offset of the section descriptor (the descriptor is stored *at the end* of the section). + #[allow(dead_code)] + pub(crate) descriptor_offset: u64, + + /// Start offset of the section data (relative to the start of the segment file). + pub(crate) data_start: u64, + + /// Length of the section data that is considered valid by `data_size`. + /// + /// Some tools may append extra bytes beyond `data_size` (e.g., after abort/restart scenarios). + pub(crate) data_len: u64, +} + +impl Ewf2Section { + /// Parse a section descriptor at an absolute file offset. + pub(crate) fn parse_at(file: &File, file_len: u64, descriptor_offset: u64) -> Result { + // The caller ensures descriptor_offset points to a valid descriptor location. + let _ = file_len; + + let mut buf = [0u8; EWF2_SECTION_DESCRIPTOR_SIZE]; + read_exact_at(file, descriptor_offset, &mut buf)?; + + let stored = u32::from_le_bytes(buf[60..64].try_into().expect("len=4")); + let calculated = adler32_rfc1950(&buf[0..60]); + if stored != calculated { + return Err(Error::Corrupt( + "EWF2 section descriptor checksum mismatch".to_string(), + )); + } + + let raw = RawEwf2SectionDescriptor::read(&mut Cursor::new(&buf[..])).map_err(|e| { + Error::Invalid(format!("invalid EWF2 section descriptor encoding: {e}")) + })?; + + let section_type = Ewf2SectionType::from_u32(raw.section_type); + let data_flags = Ewf2SectionDataFlags::new(raw.data_flags); + let previous_offset = raw.previous_offset; + let data_size = raw.data_size; + let descriptor_size = raw.descriptor_size; + let padding_size = raw.padding_size; + let data_integrity_hash = raw.data_integrity_hash; + + if descriptor_size as usize != EWF2_SECTION_DESCRIPTOR_SIZE { + return Err(Error::Invalid(format!( + "unsupported EWF2 section descriptor size: {descriptor_size}" + ))); + } + + // The descriptor links to the previous section descriptor (by offset). The section data begins + // after the previous descriptor (or after the file header for the first section). + let data_start = if previous_offset == 0 { + EWF2_FILE_HEADER_SIZE as u64 + } else { + previous_offset.saturating_add(descriptor_size as u64) + }; + + if data_start > descriptor_offset { + return Err(Error::Invalid( + "EWF2 section data start offset out of bounds".to_string(), + )); + } + + // `data_size` is the amount of data considered part of the section (including any padding + // described by `padding_size`). If tools appended extra bytes between sections, we ignore them. + let max_len = descriptor_offset.saturating_sub(data_start); + let data_len = data_size.min(max_len); + + Ok(Self { + section_type, + data_flags, + previous_offset, + data_size, + descriptor_size, + padding_size, + data_integrity_hash, + descriptor_offset, + data_start, + data_len, + }) + } + + /// Scan all section descriptors for a segment file. + pub(crate) fn scan(file: &File, file_len: u64) -> Result> { + let min_len = + (EWF2_FILE_HEADER_SIZE as u64).saturating_add(EWF2_SECTION_DESCRIPTOR_SIZE as u64); + if file_len < min_len { + return Err(Error::Invalid("file too small for EWF2".to_string())); + } + + let mut sections_rev: Vec = Vec::new(); + let mut desc_off = file_len + .checked_sub(EWF2_SECTION_DESCRIPTOR_SIZE as u64) + .ok_or_else(|| Error::Invalid("file too small for EWF2".to_string()))?; + + // Hard guard against loops on corrupt inputs. + for _ in 0..1_000_000u32 { + let section = Self::parse_at(file, file_len, desc_off)?; + let prev = section.previous_offset; + sections_rev.push(section); + + if prev == 0 { + break; + } + if prev >= desc_off { + return Err(Error::Invalid( + "EWF2 previous section offset does not move backwards".to_string(), + )); + } + desc_off = prev; + } + + if sections_rev.is_empty() { + return Err(Error::Invalid("no EWF2 sections found".to_string())); + } + + sections_rev.reverse(); + Ok(sections_rev) + } +} + +/// Writes a canonical EWF2 section descriptor to bytes. +/// +/// This is used by the writer; the reader parses the same format via [`Ewf2Section::parse_at`]. +pub(crate) fn make_ewf2_section_descriptor( + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, + previous_offset: u64, + data_size: u64, + padding_size: u32, + data_integrity_hash: [u8; 16], +) -> [u8; EWF2_SECTION_DESCRIPTOR_SIZE] { + // This is intentionally infallible: we write into an in-memory buffer of fixed size. + let desc = RawEwf2SectionDescriptor { + section_type: section_type.as_u32(), + data_flags: data_flags.raw(), + previous_offset, + data_size, + descriptor_size: EWF2_SECTION_DESCRIPTOR_SIZE as u32, + padding_size, + data_integrity_hash, + reserved: [0u8; 12], + checksum: 0, + }; + + let mut raw = [0u8; EWF2_SECTION_DESCRIPTOR_SIZE]; + desc.write(&mut Cursor::new(&mut raw[..])) + .expect("in-memory write cannot fail"); + + let checksum = adler32_rfc1950(&raw[..EWF2_SECTION_DESCRIPTOR_SIZE - 4]); + raw[EWF2_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); + raw +} + +#[cfg(test)] +mod tests { + use super::*; + use std::io::Write as _; + + #[test] + fn test_make_ewf2_section_descriptor_roundtrips_raw_fields_and_checksum() { + let section_type = Ewf2SectionType::SectorData; + let data_flags = Ewf2SectionDataFlags::MD5_HASHED; + let previous_offset = 0x1122_3344_5566_7788; + let data_size = 0x99aa_bbcc_ddee_ff00; + let padding_size = 0x1234_5678; + let data_integrity_hash: [u8; 16] = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16]; + + let bytes = make_ewf2_section_descriptor( + section_type, + data_flags, + previous_offset, + data_size, + padding_size, + data_integrity_hash, + ); + + let stored = u32::from_le_bytes(bytes[60..64].try_into().expect("len=4")); + let calculated = adler32_rfc1950(&bytes[0..60]); + assert_eq!(stored, calculated); + + let raw = + RawEwf2SectionDescriptor::read(&mut Cursor::new(&bytes[..])).expect("parse succeeds"); + assert_eq!(raw.section_type, section_type.as_u32()); + assert_eq!(raw.data_flags, data_flags.raw()); + assert_eq!(raw.previous_offset, previous_offset); + assert_eq!(raw.data_size, data_size); + assert_eq!(raw.descriptor_size, EWF2_SECTION_DESCRIPTOR_SIZE as u32); + assert_eq!(raw.padding_size, padding_size); + assert_eq!(raw.data_integrity_hash, data_integrity_hash); + assert_eq!(raw.reserved, [0u8; 12]); + assert_eq!(raw.checksum, stored); + } + + #[test] + fn test_parse_at_rejects_checksum_mismatch() { + let good = make_ewf2_section_descriptor( + Ewf2SectionType::Done, + Ewf2SectionDataFlags::new(0), + 0, + 0, + 0, + [0u8; 16], + ); + + let mut bad = good; + bad[0] ^= 0x01; + + let mut tmp = tempfile::NamedTempFile::new().expect("tempfile"); + tmp.write_all(&vec![0u8; EWF2_FILE_HEADER_SIZE]) + .expect("write header"); + let descriptor_offset = EWF2_FILE_HEADER_SIZE as u64; + tmp.write_all(&bad).expect("write descriptor"); + + let file = std::fs::File::open(tmp.path()).expect("open"); + let file_len = file.metadata().expect("metadata").len(); + + let err = Ewf2Section::parse_at(&file, file_len, descriptor_offset).unwrap_err(); + assert!(matches!(err, Error::Corrupt(msg) if msg.contains("checksum mismatch"))); + } + + #[test] + fn test_section_data_flags_preserves_unknown_bits() { + let raw = 0x8000_0000u32 | Ewf2SectionDataFlags::MD5_HASHED.bits(); + let flags = Ewf2SectionDataFlags::new(raw); + assert_eq!(flags.raw(), raw); + assert!(flags.has_md5_integrity_hash()); + assert!(!flags.is_encrypted()); + } +} diff --git a/crates/ewf/src/info.rs b/crates/ewf/src/info.rs index 75a89a6..3ad5828 100644 --- a/crates/ewf/src/info.rs +++ b/crates/ewf/src/info.rs @@ -4,6 +4,8 @@ //! etc.). This module provides a small, stable summary for callers who mostly care about //! “what kind of image is this?” and “how big is it?”. +use std::fmt; + /// High-level EWF format identifier. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum EwfFormat { @@ -19,6 +21,82 @@ pub enum EwfFormat { Lx01, } +/// libewf-compatible EWF *file format* classification. +/// +/// This corresponds to what libewf exposes via `libewf_handle_get_format()` and what +/// `ewfinfo` prints under "EWF information" → "File format". +/// +/// References (libewf): +/// - `external/libewf/ewftools/info_handle.c` (`info_handle_ewf_information_fprint`) +/// - `external/libewf/libewf/libewf_io_handle.c` (default format initialization) +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum EwfFileFormat { + /// "original EWF" + OriginalEwf, + /// "SMART" + Smart, + /// "FTK Imager" + FtkImager, + EnCase1, + EnCase2, + EnCase3, + EnCase4, + EnCase5, + EnCase6, + EnCase7, + Linen5, + Linen6, + Linen7, + /// "EWFX (extended EWF)" + Ewfx, + /// "Logical Evidence File (LEF) EnCase 5" + LogicalEnCase5, + /// "Logical Evidence File (LEF) EnCase 6" + LogicalEnCase6, + /// "Logical Evidence File (LEF) EnCase 7" + LogicalEnCase7, + /// "EnCase 7 (version 2)" + EnCase7V2, + /// "Logical Evidence File (LEF) EnCase 7 (version 2)" + LogicalEnCase7V2, + /// Unknown / not classified. + Unknown, +} + +impl EwfFileFormat { + /// Returns the exact libewf `ewfinfo` display string for this format. + pub fn as_ewfinfo_str(self) -> &'static str { + match self { + Self::OriginalEwf => "original EWF", + Self::Smart => "SMART", + Self::FtkImager => "FTK Imager", + Self::EnCase1 => "EnCase 1", + Self::EnCase2 => "EnCase 2", + Self::EnCase3 => "EnCase 3", + Self::EnCase4 => "EnCase 4", + Self::EnCase5 => "EnCase 5", + Self::EnCase6 => "EnCase 6", + Self::EnCase7 => "EnCase 7", + Self::Linen5 => "linen 5", + Self::Linen6 => "linen 6", + Self::Linen7 => "linen 7", + Self::Ewfx => "EWFX (extended EWF)", + Self::LogicalEnCase5 => "Logical Evidence File (LEF) EnCase 5", + Self::LogicalEnCase6 => "Logical Evidence File (LEF) EnCase 6", + Self::LogicalEnCase7 => "Logical Evidence File (LEF) EnCase 7", + Self::EnCase7V2 => "EnCase 7 (version 2)", + Self::LogicalEnCase7V2 => "Logical Evidence File (LEF) EnCase 7 (version 2)", + Self::Unknown => "unknown", + } + } +} + +impl fmt::Display for EwfFileFormat { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + f.write_str(self.as_ewfinfo_str()) + } +} + /// Compression algorithm identifier (not a guarantee every chunk is compressed). #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum EwfCompression { @@ -32,6 +110,8 @@ pub enum EwfCompression { #[derive(Debug, Clone, PartialEq, Eq)] pub struct EwfInfo { pub format: EwfFormat, + /// libewf-compatible file format classification. + pub file_format: EwfFileFormat, pub media_size: u64, pub chunk_size: usize, pub chunk_count: u64, diff --git a/crates/ewf/src/lib.rs b/crates/ewf/src/lib.rs index 61468fb..47e15d7 100644 --- a/crates/ewf/src/lib.rs +++ b/crates/ewf/src/lib.rs @@ -10,14 +10,18 @@ //! expand format coverage (EWF2, delta/shadow files, write resume, etc.). mod error; +mod ewf1; +mod ewf2; mod info; +mod util; pub mod delta; +pub mod metadata; pub mod reader; pub mod writer; pub use delta::EwfDelta; pub use error::{Error, Result}; -pub use info::{EwfCompression, EwfFormat, EwfInfo}; +pub use info::{EwfCompression, EwfFileFormat, EwfFormat, EwfInfo}; pub use reader::{EwfReader, LefEntry, LefExtent, LefReader, VerifyOptions}; pub use writer::{Ewf2CompressionMethod, Ewf2Writer, Ewf2WriterOptions, EwfWriter}; diff --git a/crates/ewf/src/metadata.rs b/crates/ewf/src/metadata.rs new file mode 100644 index 0000000..b8188c1 --- /dev/null +++ b/crates/ewf/src/metadata.rs @@ -0,0 +1,100 @@ +//! Spec-oriented metadata extracted from EWF images. +//! +//! This module is intentionally **format/spec** focused. It exposes structured metadata that can be +//! consumed by binaries (like `ewfinfo`) or other library users without coupling the `ewf` crate to +//! libewf’s CLI/reporting surface area. +//! +//! In particular: +//! - Types here model *what is in the image* (or what can be inferred from it). +//! - Formatting and presentation (labels, wrapping, date formatting, etc.) belongs to binaries. + +use std::path::PathBuf; + +use crate::{EwfCompression, EwfFileFormat, EwfFormat}; + +/// A contiguous run expressed in logical sectors. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct SectorRun { + /// Start sector index (LBA). + pub start_sector: u64, + /// Number of sectors in the run. + pub sector_count: u64, +} + +/// High-level media type. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum MediaType { + RemovableDisk, + FixedDisk, + OpticalDisk, + SingleFiles, + MemoryRam, + Unknown, +} + +/// A coarse compression “level” indicator used by some EWF1 variants. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum CompressionLevel { + NoCompression, + GoodFastCompression, + BestCompression, + Unknown, + NotRecorded, +} + +/// Header (acquiry) values extracted from EWF1 header sections. +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub struct HeaderValues { + pub case_number: Option, + pub evidence_number: Option, + pub description: Option, + pub examiner_name: Option, + pub notes: Option, + pub acquisition_datetime: Option, + pub system_datetime: Option, + pub acquisition_os: Option, + pub acquisition_software: Option, + pub acquisition_software_version: Option, + /// A password/hash value if present. Absence means “not set”. + pub password: Option, +} + +/// Image-level digests (when present). +#[derive(Debug, Clone, Default, PartialEq, Eq)] +pub struct ImageDigests { + pub md5: Option<[u8; 16]>, + pub sha1: Option<[u8; 20]>, +} + +/// Metadata extracted from an EWF image set. +#[derive(Debug, Clone, PartialEq, Eq)] +pub struct ImageMetadata { + pub format: EwfFormat, + /// libewf-compatible file format classification. + pub file_format: EwfFileFormat, + pub segment_paths: Vec, + + /// Segment file version (EWF2/EVF2 only). + pub segment_file_version: Option<(u16, u16)>, + + pub sectors_per_chunk: u32, + /// The number of sectors used as error granularity (0 if not recorded). + pub error_granularity: u32, + pub bytes_per_sector: u32, + pub number_of_sectors: u64, + pub media_size: u64, + + pub media_type: MediaType, + pub is_physical: bool, + + pub compression_method: EwfCompression, + pub compression_level: CompressionLevel, + pub set_identifier: Option<[u8; 16]>, + + pub header_values: HeaderValues, + pub digests: ImageDigests, + + pub sessions: Vec, + pub tracks: Vec, + pub acquisition_read_errors: Vec, +} diff --git a/crates/ewf/src/reader.rs b/crates/ewf/src/reader.rs index 3f8b6b2..53c65f2 100644 --- a/crates/ewf/src/reader.rs +++ b/crates/ewf/src/reader.rs @@ -9,7 +9,19 @@ //! //! EWF2 (`.Ex01`, `.Lx01`) and logical evidence (`.L01`) are wired in later in this task series. -use crate::{Error, EwfCompression, EwfFormat, EwfInfo, Result}; +use crate::ewf1::runs::{Ewf1Error2Section, Ewf1SessionSection}; +use crate::ewf1::section::{Ewf1SectionDescriptor, Ewf1SectionType}; +use crate::ewf1::volume as ewf1_volume; +use crate::ewf1::{ + EWF1_EVF_SIGNATURE, EWF1_FILE_HEADER_SIZE, EWF1_LVF_SIGNATURE, EWF1_TABLE_HEADER_SIZE, +}; +use crate::ewf2::chunk::Ewf2ChunkDataFlags; +use crate::ewf2::file_header::{ + EWF2_EVF_SIGNATURE, EWF2_FILE_HEADER_SIZE, EWF2_LEF_SIGNATURE, Ewf2FileHeader, Ewf2Kind, +}; +use crate::ewf2::section::{Ewf2Section, Ewf2SectionType}; +use crate::util::{adler32_rfc1950, read_exact_at, read_file_range}; +use crate::{Error, EwfCompression, EwfFileFormat, EwfFormat, EwfInfo, Result}; use flate2::read::ZlibDecoder; use lru::LruCache; use md5::{Digest as _, Md5}; @@ -20,23 +32,45 @@ use std::num::NonZeroUsize; use std::path::{Path, PathBuf}; use std::sync::Mutex; -#[cfg(unix)] -use std::os::unix::fs::FileExt as _; - -#[cfg(windows)] -use std::os::windows::fs::FileExt as _; - // --- EWF signatures (file header; first 8 bytes) --- -const EWF1_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; // "EVF\t\r\n\xff\0" -const EWF1_LVF_SIGNATURE: [u8; 8] = [0x4c, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; // "LVF\t\r\n\xff\0" (logical evidence) -const EWF2_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x32, 0x0d, 0x0a, 0x81, 0x00]; // "EVF2\r\n\x81\0" -const EWF2_LEF_SIGNATURE: [u8; 8] = [0x4c, 0x45, 0x46, 0x32, 0x0d, 0x0a, 0x81, 0x00]; // "LEF2\r\n\x81\0" -const ADCRYPT_SIGNATURE: [u8; 8] = [0x41, 0x44, 0x43, 0x52, 0x59, 0x50, 0x54, 0x00]; // "ADCRYPT\0" +const ADCRYPT_SIGNATURE: [u8; 8] = *b"ADCRYPT\0"; + +fn detect_ewf1_file_format(container: EwfFormat, volume_section_body: &[u8]) -> EwfFileFormat { + // libewf's "file format" classification is broader than our container format. + // + // We keep this mapping intentionally conservative and spec-driven: + // - SMART can be detected from the 94-byte volume section signature ("SMART"). + // - The original EWF (ASR02) "EWF specification" volume section uses a 5-byte signature that + // contains the *prefix* of the EWF file header signature (see EWF spec: volume section). + // - EnCase/FTK/linen EWF-E01 all share the 1052-byte volume section variant and cannot be + // reliably distinguished without additional heuristics, so we match libewf's default: + // "EnCase 6". + + // The container naming scheme already distinguishes `.s01` images in our reader. + if matches!(container, EwfFormat::S01) { + return EwfFileFormat::Smart; + } + + // EWF spec "volume section" (94-byte variant): + // - offset 85, size 5: signature + // - "SMART" for SMART + // - EWF file header signature prefix for original EWF + if volume_section_body.len() >= 94 && volume_section_body.len() < 1052 { + let sig = &volume_section_body[85..90]; + if sig == b"SMART" { + return EwfFileFormat::Smart; + } + // Prefix of "EVF\x09\x0d\x0a\xff\x00" (first 5 bytes). + const EWF_SIG_PREFIX: [u8; 5] = [0x45, 0x56, 0x46, 0x09, 0x0d]; + if sig == EWF_SIG_PREFIX { + return EwfFileFormat::OriginalEwf; + } + return EwfFileFormat::Unknown; + } -// --- EWF1 constants --- -const EWF1_FILE_HEADER_SIZE: usize = 8 + 1 + 2 + 2; // 13 -const EWF1_SECTION_DESCRIPTOR_SIZE: usize = 16 + 8 + 8 + 40 + 4; // 76 -const EWF1_TABLE_HEADER_SIZE: usize = 4 + 4 + 8 + 4 + 4; // 24 + // Default for EWF-E01. + EwfFileFormat::EnCase6 +} /// Random-access reader over an EWF image set. /// @@ -154,75 +188,14 @@ struct Ewf1ChunkGroup { } // --- EWF2 constants --- -const EWF2_FILE_HEADER_SIZE: usize = 32; -const EWF2_SECTION_DESCRIPTOR_SIZE: usize = 64; const EWF2_TABLE_HEADER_SIZE: usize = 32; // 20 bytes header + 12 bytes alignment padding const EWF2_TABLE_ENTRY_SIZE: usize = 16; const EWF2_TABLE_FOOTER_SIZE: usize = 16; // 4 bytes footer + 12 bytes alignment padding -const EWF2_SECTION_TYPE_DEVICE_INFORMATION: u32 = 0x0000_0001; -const EWF2_SECTION_TYPE_CASE_DATA: u32 = 0x0000_0002; -const EWF2_SECTION_TYPE_SECTOR_DATA: u32 = 0x0000_0003; -const EWF2_SECTION_TYPE_SECTOR_TABLE: u32 = 0x0000_0004; -const EWF2_SECTION_TYPE_NEXT: u32 = 0x0000_000d; -const EWF2_SECTION_TYPE_DONE: u32 = 0x0000_000f; -const EWF2_SECTION_TYPE_SINGLE_FILES_DATA: u32 = 0x0000_0020; - -const EWF2_SECTION_DATA_FLAG_MD5HASHED: u32 = 0x0000_0001; -const EWF2_SECTION_DATA_FLAG_ENCRYPTED: u32 = 0x0000_0002; - -const EWF2_CHUNK_DATA_FLAG_COMPRESSED: u32 = 0x0000_0001; -const EWF2_CHUNK_DATA_FLAG_CHECKSUMED: u32 = 0x0000_0002; -const EWF2_CHUNK_DATA_FLAG_PATTERNFILL: u32 = 0x0000_0004; - const EWF2_COMPRESSION_NONE: u16 = 0; const EWF2_COMPRESSION_LZ: u16 = 1; const EWF2_COMPRESSION_BZIP2: u16 = 2; -#[derive(Debug, Clone, Copy, PartialEq, Eq)] -enum Ewf2Kind { - Ex01, - Lx01, -} - -#[derive(Debug)] -struct Ewf2FileHeader { - kind: Ewf2Kind, - #[allow(dead_code)] - major: u8, - #[allow(dead_code)] - minor: u8, - compression_method: u16, - segment_number: u32, - set_id: [u8; 16], -} - -#[derive(Debug, Clone)] -struct Ewf2Section { - section_type: u32, - data_flags: u32, - previous_offset: u64, - #[allow(dead_code)] - data_size: u64, - #[allow(dead_code)] - descriptor_size: u32, - padding_size: u32, - #[allow(dead_code)] - data_integrity_hash: [u8; 16], - - /// Offset of the section descriptor (the descriptor is stored *at the end* of the section). - #[allow(dead_code)] - descriptor_offset: u64, - - /// Start offset of the section data (relative to the start of the segment file). - data_start: u64, - - /// Length of the section data that is considered valid by `data_size`. - /// - /// Some tools may append extra bytes beyond `data_size` (e.g., after abort/restart scenarios). - data_len: u64, -} - #[derive(Debug)] struct Ewf2Segment { #[allow(dead_code)] @@ -238,7 +211,7 @@ struct Ewf2Segment { struct Ewf2TableEntry { offset_raw: [u8; 8], size: u32, - flags: u32, + flags: Ewf2ChunkDataFlags, } #[derive(Debug)] @@ -340,6 +313,17 @@ impl EwfReader { } } + /// Extracts spec-oriented image metadata. + /// + /// This is intended to be consumed by binaries (like `ewfinfo`) that own the presentation + /// layer (labels, date formatting, wrapping, etc.). + pub fn image_metadata(&self) -> Result { + match &self.inner { + InnerReader::Ewf1(r) => r.image_metadata(), + InnerReader::Ewf2(r) => r.image_metadata(), + } + } + /// Verifies the image set according to `opts`. /// /// By default this verifies chunk decoding (it will read and decode every chunk). @@ -433,11 +417,8 @@ impl Ewf2Reader { set_id = Some(hdr.set_id); } - let sections = parse_ewf2_section_descriptors(&file, file_len)?; - if sections - .iter() - .any(|s| (s.data_flags & EWF2_SECTION_DATA_FLAG_ENCRYPTED) != 0) - { + let sections = Ewf2Section::scan(&file, file_len)?; + if sections.iter().any(|s| s.data_flags.is_encrypted()) { return Err(Error::Unsupported( "EWF2 encrypted images are not supported (matches libewf)".to_string(), )); @@ -462,12 +443,12 @@ impl Ewf2Reader { return Err(Error::Invalid("EWF2 segment has no sections".to_string())); }; if is_last { - if last.section_type != EWF2_SECTION_TYPE_DONE { + if last.section_type != Ewf2SectionType::Done { return Err(Error::Invalid( "EWF2 last segment does not end with done section".to_string(), )); } - } else if last.section_type != EWF2_SECTION_TYPE_NEXT { + } else if last.section_type != Ewf2SectionType::Next { return Err(Error::Invalid( "EWF2 non-last segment does not end with next section".to_string(), )); @@ -479,7 +460,7 @@ impl Ewf2Reader { let any_sector_data = segments.iter().any(|seg| { seg.sections .iter() - .any(|s| s.section_type == EWF2_SECTION_TYPE_SECTOR_DATA) + .any(|s| s.section_type == Ewf2SectionType::SectorData) }); if !any_sector_data { return Err(Error::Invalid( @@ -495,12 +476,12 @@ impl Ewf2Reader { let device_section = first .sections .iter() - .find(|s| s.section_type == EWF2_SECTION_TYPE_DEVICE_INFORMATION) + .find(|s| s.section_type == Ewf2SectionType::DeviceInformation) .ok_or_else(|| Error::Invalid("missing EWF2 device information section".to_string()))?; let case_section = first .sections .iter() - .find(|s| s.section_type == EWF2_SECTION_TYPE_CASE_DATA) + .find(|s| s.section_type == Ewf2SectionType::CaseData) .ok_or_else(|| Error::Invalid("missing EWF2 case data section".to_string()))?; let device_tags = parse_ewf2_main_object_tags(&read_ewf2_compressed_object_string( @@ -552,7 +533,7 @@ impl Ewf2Reader { let mut groups: Vec = Vec::new(); for (seg_idx, seg) in segments.iter().enumerate() { for section in &seg.sections { - if section.section_type == EWF2_SECTION_TYPE_SECTOR_TABLE { + if section.section_type == Ewf2SectionType::SectorTable { let (first_chunk, entries) = parse_ewf2_sector_table_section(seg, section)?; groups.push(Ewf2ChunkGroup { segment_idx: seg_idx, @@ -629,6 +610,7 @@ impl Ewf2Reader { EwfInfo { format: self.format(), + file_format: EwfFileFormat::EnCase7V2, media_size: self.media_size, chunk_size: self.chunk_size, chunk_count: self.chunk_count, @@ -637,6 +619,153 @@ impl Ewf2Reader { } } + fn image_metadata(&self) -> Result { + use crate::metadata::{ + CompressionLevel, HeaderValues, ImageDigests, ImageMetadata, MediaType, + }; + + let first = self + .segments + .first() + .ok_or_else(|| Error::Invalid("missing first EWF2 segment".to_string()))?; + + // Parse file header to extract set-id and version. + let hdr = parse_ewf2_file_header(&first.file)?; + + let device_section = first + .sections + .iter() + .find(|s| s.section_type == Ewf2SectionType::DeviceInformation) + .ok_or_else(|| Error::Invalid("missing EWF2 device information section".to_string()))?; + let case_section = first + .sections + .iter() + .find(|s| s.section_type == Ewf2SectionType::CaseData) + .ok_or_else(|| Error::Invalid("missing EWF2 case data section".to_string()))?; + + let device_tags = parse_ewf2_main_object_tags(&read_ewf2_compressed_object_string( + first, + device_section, + self.compression_method, + )?)?; + let case_tags = parse_ewf2_main_object_tags(&read_ewf2_compressed_object_string( + first, + case_section, + self.compression_method, + )?)?; + + let bytes_per_sector = parse_tag_u32(&device_tags, "bp")?; + let number_of_sectors = parse_tag_u64(&device_tags, "ts")?; + let sectors_per_chunk = parse_tag_u32(&case_tags, "sb")?; + let error_granularity = match case_tags.get("gr") { + Some(v) => v + .trim() + .parse::() + .map_err(|_| Error::Invalid(format!("invalid EWF2 `gr` value: `{v}`")))?, + None => 0, + }; + + let media_size = number_of_sectors + .checked_mul(bytes_per_sector as u64) + .ok_or_else(|| Error::Invalid("media size overflow".to_string()))?; + + let media_type = match device_tags.get("dt").map(|s| s.trim()) { + Some("r") => MediaType::RemovableDisk, + Some("f") => MediaType::FixedDisk, + Some("o") => MediaType::OpticalDisk, + Some("m") => MediaType::MemoryRam, + _ => MediaType::FixedDisk, + }; + + let is_physical = device_tags + .get("ph") + .and_then(|s| s.trim().parse::().ok()) + .map(|v| v != 0) + .unwrap_or(false); + + // Hash sections: scan the segment set for MD5/SHA1 hash sections. + let mut md5: Option<[u8; 16]> = None; + let mut sha1: Option<[u8; 20]> = None; + + for seg in &self.segments { + for section in &seg.sections { + let mut raw_len = section.data_len; + if (section.padding_size as u64) <= raw_len { + raw_len = raw_len.saturating_sub(section.padding_size as u64); + } + if raw_len == 0 { + continue; + } + + if section.section_type == Ewf2SectionType::Md5Hash { + let bytes = read_file_range( + &seg.file, + seg.file_len, + section.data_start, + section.data_start.saturating_add(raw_len), + )?; + if bytes.len() >= 16 { + let mut m = [0u8; 16]; + m.copy_from_slice(&bytes[0..16]); + md5 = Some(m); + } + } else if section.section_type == Ewf2SectionType::Sha1Hash { + let bytes = read_file_range( + &seg.file, + seg.file_len, + section.data_start, + section.data_start.saturating_add(raw_len), + )?; + if bytes.len() >= 20 { + let mut s = [0u8; 20]; + s.copy_from_slice(&bytes[0..20]); + sha1 = Some(s); + } + } + } + } + + let compression_method = match hdr.compression_method { + EWF2_COMPRESSION_NONE => EwfCompression::None, + EWF2_COMPRESSION_LZ => EwfCompression::Zlib, + EWF2_COMPRESSION_BZIP2 => EwfCompression::Bzip2, + other => EwfCompression::Unknown(other), + }; + + let compression_level = match hdr.compression_method { + EWF2_COMPRESSION_NONE => CompressionLevel::NoCompression, + _ => CompressionLevel::Unknown, + }; + + let set_identifier = if hdr.set_id.iter().any(|&b| b != 0) { + Some(hdr.set_id) + } else { + None + }; + + Ok(ImageMetadata { + format: self.format(), + file_format: EwfFileFormat::EnCase7V2, + segment_paths: self.segments.iter().map(|s| s.path.clone()).collect(), + segment_file_version: Some((hdr.major as u16, hdr.minor as u16)), + sectors_per_chunk, + error_granularity, + bytes_per_sector, + number_of_sectors, + media_size, + media_type, + is_physical, + compression_method, + compression_level, + set_identifier, + header_values: HeaderValues::default(), + digests: ImageDigests { md5, sha1 }, + sessions: Vec::new(), + tracks: Vec::new(), + acquisition_read_errors: Vec::new(), + }) + } + fn verify(&self, opts: VerifyOptions) -> Result<()> { if opts.verify_section_md5 { self.verify_section_md5(opts.verify_sector_data_section_md5)?; @@ -653,11 +782,11 @@ impl Ewf2Reader { fn verify_section_md5(&self, include_sector_data: bool) -> Result<()> { for seg in &self.segments { for section in &seg.sections { - let has_md5 = (section.data_flags & EWF2_SECTION_DATA_FLAG_MD5HASHED) != 0; + let has_md5 = section.data_flags.has_md5_integrity_hash(); if !has_md5 { continue; } - if !include_sector_data && section.section_type == EWF2_SECTION_TYPE_SECTOR_DATA { + if !include_sector_data && section.section_type == Ewf2SectionType::SectorData { continue; } let digest = md5_file_range( @@ -669,7 +798,7 @@ impl Ewf2Reader { if digest != section.data_integrity_hash { return Err(Error::Corrupt(format!( "EWF2 section MD5 mismatch (type=0x{:08x})", - section.section_type + section.section_type.as_u32() ))); } } @@ -715,10 +844,10 @@ impl Ewf2Reader { Error::Invalid("chunk index out of range in sector table".to_string()) })?; - let is_compressed = (entry.flags & EWF2_CHUNK_DATA_FLAG_COMPRESSED) != 0; - let is_checksumed = (entry.flags & EWF2_CHUNK_DATA_FLAG_CHECKSUMED) != 0; + let is_compressed = entry.flags.contains(Ewf2ChunkDataFlags::COMPRESSED); + let is_checksumed = entry.flags.contains(Ewf2ChunkDataFlags::CHECKSUMED); let is_pattern_fill = - is_compressed && (entry.flags & EWF2_CHUNK_DATA_FLAG_PATTERNFILL) != 0; + is_compressed && entry.flags.contains(Ewf2ChunkDataFlags::PATTERNFILL); let mut out = vec![0u8; self.chunk_size]; @@ -840,141 +969,7 @@ impl Ewf2Reader { fn parse_ewf2_file_header(file: &File) -> Result { let mut buf = [0u8; EWF2_FILE_HEADER_SIZE]; read_exact_at(file, 0, &mut buf)?; - - let signature: [u8; 8] = buf[0..8].try_into().expect("len=8"); - let kind = if signature == EWF2_EVF_SIGNATURE { - Ewf2Kind::Ex01 - } else if signature == EWF2_LEF_SIGNATURE { - Ewf2Kind::Lx01 - } else { - return Err(Error::Invalid("not an EWF2 segment file".to_string())); - }; - - let major = buf[8]; - let minor = buf[9]; - if major != 2 { - return Err(Error::Invalid(format!( - "unsupported EWF2 major version: {major}" - ))); - } - - let compression_method = u16::from_le_bytes(buf[10..12].try_into().expect("len=2")); - let segment_number = u32::from_le_bytes(buf[12..16].try_into().expect("len=4")); - let set_id: [u8; 16] = buf[16..32].try_into().expect("len=16"); - - Ok(Ewf2FileHeader { - kind, - major, - minor, - compression_method, - segment_number, - set_id, - }) -} - -fn parse_ewf2_section_descriptors(file: &File, file_len: u64) -> Result> { - let min_len = - (EWF2_FILE_HEADER_SIZE as u64).saturating_add(EWF2_SECTION_DESCRIPTOR_SIZE as u64); - if file_len < min_len { - return Err(Error::Invalid("file too small for EWF2".to_string())); - } - - let mut sections_rev: Vec = Vec::new(); - let mut desc_off = file_len - .checked_sub(EWF2_SECTION_DESCRIPTOR_SIZE as u64) - .ok_or_else(|| Error::Invalid("file too small for EWF2".to_string()))?; - - // Hard guard against loops on corrupt inputs. - for _ in 0..1_000_000u32 { - let section = parse_ewf2_section_descriptor_at(file, file_len, desc_off)?; - let prev = section.previous_offset; - sections_rev.push(section); - - if prev == 0 { - break; - } - if prev >= desc_off { - return Err(Error::Invalid( - "EWF2 previous section offset does not move backwards".to_string(), - )); - } - desc_off = prev; - } - - if sections_rev.is_empty() { - return Err(Error::Invalid("no EWF2 sections found".to_string())); - } - - sections_rev.reverse(); - Ok(sections_rev) -} - -fn parse_ewf2_section_descriptor_at( - file: &File, - file_len: u64, - descriptor_offset: u64, -) -> Result { - let _ = file_len; - let mut buf = [0u8; EWF2_SECTION_DESCRIPTOR_SIZE]; - read_exact_at(file, descriptor_offset, &mut buf)?; - - let stored = u32::from_le_bytes(buf[60..64].try_into().expect("len=4")); - let calculated = adler32_rfc1950(&buf[0..60]); - if stored != calculated { - return Err(Error::Corrupt( - "EWF2 section descriptor checksum mismatch".to_string(), - )); - } - - let section_type = u32::from_le_bytes(buf[0..4].try_into().expect("len=4")); - let data_flags = u32::from_le_bytes(buf[4..8].try_into().expect("len=4")); - // If the section has an MD5 integrity hash, the hash covers the (possibly encrypted) section data. - // We currently do not verify it during open; chunk-level checksums and section descriptor checksums - // provide basic corruption detection already. - let _has_md5_integrity_hash = (data_flags & EWF2_SECTION_DATA_FLAG_MD5HASHED) != 0; - let previous_offset = u64::from_le_bytes(buf[8..16].try_into().expect("len=8")); - let data_size = u64::from_le_bytes(buf[16..24].try_into().expect("len=8")); - let descriptor_size = u32::from_le_bytes(buf[24..28].try_into().expect("len=4")); - let padding_size = u32::from_le_bytes(buf[28..32].try_into().expect("len=4")); - let data_integrity_hash: [u8; 16] = buf[32..48].try_into().expect("len=16"); - - if descriptor_size as usize != EWF2_SECTION_DESCRIPTOR_SIZE { - return Err(Error::Invalid(format!( - "unsupported EWF2 section descriptor size: {descriptor_size}" - ))); - } - - // The descriptor links to the previous section descriptor (by offset). The section data begins - // after the previous descriptor (or after the file header for the first section). - let data_start = if previous_offset == 0 { - EWF2_FILE_HEADER_SIZE as u64 - } else { - previous_offset.saturating_add(descriptor_size as u64) - }; - - if data_start > descriptor_offset { - return Err(Error::Invalid( - "EWF2 section data start offset out of bounds".to_string(), - )); - } - - // `data_size` is the amount of data considered part of the section (including any padding - // described by `padding_size`). If tools appended extra bytes between sections, we ignore them. - let max_len = descriptor_offset.saturating_sub(data_start); - let data_len = data_size.min(max_len); - - Ok(Ewf2Section { - section_type, - data_flags, - previous_offset, - data_size, - descriptor_size, - padding_size, - data_integrity_hash, - descriptor_offset, - data_start, - data_len, - }) + Ewf2FileHeader::parse(&buf) } fn read_ewf2_compressed_object_string( @@ -982,7 +977,7 @@ fn read_ewf2_compressed_object_string( section: &Ewf2Section, compression_method: u16, ) -> Result { - if (section.data_flags & EWF2_SECTION_DATA_FLAG_ENCRYPTED) != 0 { + if section.data_flags.is_encrypted() { return Err(Error::Unsupported( "EWF2 encrypted metadata sections are not supported yet".to_string(), )); @@ -1134,7 +1129,7 @@ fn parse_ewf2_sector_table_section( segment: &Ewf2Segment, section: &Ewf2Section, ) -> Result<(u64, Vec)> { - if (section.data_flags & EWF2_SECTION_DATA_FLAG_ENCRYPTED) != 0 { + if section.data_flags.is_encrypted() { return Err(Error::Unsupported( "EWF2 encrypted sector tables are not supported yet".to_string(), )); @@ -1198,7 +1193,8 @@ fn parse_ewf2_sector_table_section( for chunk in entries_bytes.chunks_exact(EWF2_TABLE_ENTRY_SIZE) { let offset_raw: [u8; 8] = chunk[0..8].try_into().expect("len=8"); let size = u32::from_le_bytes(chunk[8..12].try_into().expect("len=4")); - let flags = u32::from_le_bytes(chunk[12..16].try_into().expect("len=4")); + let flags_raw = u32::from_le_bytes(chunk[12..16].try_into().expect("len=4")); + let flags = Ewf2ChunkDataFlags::from_raw(flags_raw); entries.push(Ewf2TableEntry { offset_raw, size, @@ -1268,6 +1264,11 @@ impl Ewf1Reader { fn info(&self) -> EwfInfo { EwfInfo { format: self.format(), + file_format: match self.format() { + EwfFormat::S01 => EwfFileFormat::Smart, + // For EWF-E01, libewf defaults to "EnCase 6". + _ => EwfFileFormat::EnCase6, + }, media_size: self.media_size, chunk_size: self.chunk_size, chunk_count: self.chunk_count, @@ -1276,6 +1277,209 @@ impl Ewf1Reader { } } + fn image_metadata(&self) -> Result { + use crate::metadata::{ + CompressionLevel, HeaderValues, ImageDigests, ImageMetadata, MediaType, + }; + + let first = self + .segments + .first() + .ok_or_else(|| Error::Invalid("missing first segment".to_string()))?; + let last = self + .segments + .last() + .ok_or_else(|| Error::Invalid("missing last segment".to_string()))?; + + let first_sections = + Ewf1SectionDescriptor::scan(&first.file, first.file_len, EWF1_FILE_HEADER_SIZE as u64)?; + let last_sections = + Ewf1SectionDescriptor::scan(&last.file, last.file_len, EWF1_FILE_HEADER_SIZE as u64)?; + + let header_desc = first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Header)); + let header2_desc = first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Header2)); + let xheader_desc = first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::XHeader)); + + let volume_desc = first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Volume)) + .or_else(|| { + first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Disk)) + }) + .or_else(|| { + first_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Data)) + }) + .ok_or_else(|| { + Error::Invalid("missing required section `volume` (or `disk`/`data`)".to_string()) + })?; + + let volume_data = { + let (start, end) = volume_desc.data_range()?; + read_file_range(&first.file, first.file_len, start, end)? + }; + let file_format = detect_ewf1_file_format(self.format(), &volume_data); + let volume = + ewf1_volume::Ewf1VolumeInfo::parse_from_volume_like_section_body(&volume_data)?; + + let mut header_values = HeaderValues::default(); + + if let Some(desc) = header_desc { + let compressed = { + let (start, end) = desc.data_range()?; + read_file_range(&first.file, first.file_len, start, end)? + }; + let decompressed = zlib_decompress_all(&compressed)?; + crate::ewf1::header::Ewf1HeaderAscii::new(&decompressed) + .parse_into(&mut header_values)?; + } + + if let Some(desc) = header2_desc { + let compressed = { + let (start, end) = desc.data_range()?; + read_file_range(&first.file, first.file_len, start, end)? + }; + let decompressed = zlib_decompress_all(&compressed)?; + crate::ewf1::header::Ewf1Header2Utf16Le::new(&decompressed) + .parse_into(&mut header_values)?; + } + + // EWF-X (EWFX) optional XML header section (`xheader`), if present. + // This is where libewf stores `acquiry_software` and `acquiry_software_version`. + if let Some(desc) = xheader_desc { + let compressed = { + let (start, end) = desc.data_range()?; + read_file_range(&first.file, first.file_len, start, end)? + }; + let decompressed = zlib_decompress_all(&compressed)?; + crate::ewf1::header::EwfxXHeaderXml::new(&decompressed) + .parse_into(&mut header_values)?; + } + + crate::ewf1::header::normalize_header_values(&mut header_values); + + // Digest/hash sections live (typically) in the last segment. + let digest_desc = last_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Digest)); + let hash_desc = last_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Hash)); + + let mut md5: Option<[u8; 16]> = None; + let mut sha1: Option<[u8; 20]> = None; + + if let Some(desc) = digest_desc { + let digest_data = { + let (start, end) = desc.data_range()?; + read_file_range(&last.file, last.file_len, start, end)? + }; + if digest_data.len() >= 36 { + let mut m = [0u8; 16]; + m.copy_from_slice(&digest_data[0..16]); + md5 = Some(m); + let mut s = [0u8; 20]; + s.copy_from_slice(&digest_data[16..36]); + sha1 = Some(s); + } + } else if let Some(desc) = hash_desc { + let hash_data = { + let (start, end) = desc.data_range()?; + read_file_range(&last.file, last.file_len, start, end)? + }; + if hash_data.len() >= 16 { + let mut m = [0u8; 16]; + m.copy_from_slice(&hash_data[0..16]); + md5 = Some(m); + } + } + + // Session and error sections are optional and only show for optical media / read errors. + let sessions = if let Some(desc) = last_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Session)) + { + let data = { + let (start, end) = desc.data_range()?; + read_file_range(&last.file, last.file_len, start, end)? + }; + Ewf1SessionSection::new(&data).runs(volume.number_of_sectors)? + } else { + Vec::new() + }; + + let acquisition_read_errors = if let Some(desc) = last_sections + .iter() + .find(|s| matches!(s.section_type, Ewf1SectionType::Error2)) + { + let data = { + let (start, end) = desc.data_range()?; + read_file_range(&last.file, last.file_len, start, end)? + }; + Ewf1Error2Section::new(&data).runs()? + } else { + Vec::new() + }; + + // NOTE: libewf exposes both “sessions” and “tracks”. The EWF spec documents only the + // `session` section; for now we mirror it into tracks as well. + let tracks = sessions.clone(); + + let media_type = match volume.media_type { + ewf1_volume::Ewf1MediaType::RemovableDisk => MediaType::RemovableDisk, + ewf1_volume::Ewf1MediaType::FixedDisk => MediaType::FixedDisk, + ewf1_volume::Ewf1MediaType::OpticalDisk => MediaType::OpticalDisk, + ewf1_volume::Ewf1MediaType::SingleFiles => MediaType::SingleFiles, + ewf1_volume::Ewf1MediaType::MemoryRam => MediaType::MemoryRam, + ewf1_volume::Ewf1MediaType::Unknown(_) => MediaType::Unknown, + }; + + let compression_level = match volume.compression_level { + ewf1_volume::Ewf1VolumeCompressionLevel::NoCompression => { + CompressionLevel::NoCompression + } + ewf1_volume::Ewf1VolumeCompressionLevel::GoodFastCompression => { + CompressionLevel::GoodFastCompression + } + ewf1_volume::Ewf1VolumeCompressionLevel::BestCompression => { + CompressionLevel::BestCompression + } + ewf1_volume::Ewf1VolumeCompressionLevel::Unknown(_) => CompressionLevel::Unknown, + ewf1_volume::Ewf1VolumeCompressionLevel::NotRecorded => CompressionLevel::NotRecorded, + }; + + Ok(ImageMetadata { + format: self.format(), + file_format, + segment_paths: self.segments.iter().map(|s| s.path.clone()).collect(), + segment_file_version: None, + sectors_per_chunk: volume.sectors_per_chunk, + error_granularity: volume.error_granularity, + bytes_per_sector: volume.bytes_per_sector, + number_of_sectors: volume.number_of_sectors, + media_size: volume.media_size, + media_type, + is_physical: volume.is_physical, + compression_method: EwfCompression::Zlib, + compression_level, + set_identifier: volume.set_identifier, + header_values, + digests: ImageDigests { md5, sha1 }, + sessions, + tracks, + acquisition_read_errors, + }) + } + fn verify(&self, opts: VerifyOptions) -> Result<()> { if opts.verify_section_md5 || opts.verify_sector_data_section_md5 { // EWF1 does not define section MD5 integrity hashes the way EWF2 does. @@ -1365,8 +1569,10 @@ impl Ewf1Reader { } let sections = - parse_ewf1_section_descriptors(&file, file_len, hdr.sections_start_offset())?; - any_table2 |= sections.iter().any(|s| s.type_string == "table2"); + Ewf1SectionDescriptor::scan(&file, file_len, hdr.sections_start_offset())?; + any_table2 |= sections + .iter() + .any(|s| matches!(s.section_type, Ewf1SectionType::Table2)); parsed.push(Ewf1SegmentParsed { path: seg_path.clone(), @@ -1585,108 +1791,6 @@ impl Ewf1FileHeader { } } -#[derive(Debug, Clone)] -struct Ewf1SectionDescriptor { - start_offset: u64, - type_string: String, - size: u64, -} - -impl Ewf1SectionDescriptor { - fn parse_at(file: &File, file_len: u64, start_offset: u64) -> Result { - if start_offset >= file_len { - return Err(io::Error::from(io::ErrorKind::UnexpectedEof).into()); - } - - let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; - read_exact_at(file, start_offset, &mut raw)?; - - let stored_checksum = u32::from_le_bytes( - raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..] - .try_into() - .expect("len=4"), - ); - let calculated_checksum = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); - if stored_checksum != calculated_checksum { - return Err(Error::Corrupt( - "section descriptor checksum mismatch".to_string(), - )); - } - - let type_string = parse_ascii_nul_terminated(&raw[0..16]); - let next_offset = u64::from_le_bytes(raw[16..24].try_into().expect("len=8")); - let mut size = u64::from_le_bytes(raw[24..32].try_into().expect("len=8")); - - // libewf behavior: some writers leave size = 0, but set next_offset; infer size from that. - if size == 0 && next_offset != start_offset && next_offset >= start_offset { - size = next_offset - start_offset; - } - - Ok(Self { - start_offset, - type_string, - size, - }) - } - - fn data_range(&self) -> Result<(u64, u64)> { - let start = self - .start_offset - .checked_add(EWF1_SECTION_DESCRIPTOR_SIZE as u64) - .ok_or_else(|| Error::Invalid("section range overflow".to_string()))?; - let end = self - .start_offset - .checked_add(self.size) - .ok_or_else(|| Error::Invalid("section range overflow".to_string()))?; - Ok((start, end)) - } -} - -fn parse_ewf1_section_descriptors( - file: &File, - file_len: u64, - first_section_offset: u64, -) -> Result> { - let mut sections = Vec::new(); - let mut offset = first_section_offset; - - // Hard safety cap: avoid pathological scans on corrupted inputs. - for _ in 0..100_000 { - if offset == 0 || offset >= file_len { - break; - } - - let desc = Ewf1SectionDescriptor::parse_at(file, file_len, offset)?; - let is_last = desc.type_string == "next" || desc.type_string == "done"; - - let advance = if desc.size != 0 { - desc.size - } else { - // libewf: for last sections (`next`/`done`) some writers set size=0; advance by descriptor size. - EWF1_SECTION_DESCRIPTOR_SIZE as u64 - }; - - if advance == 0 { - return Err(Error::Invalid( - "zero advance while scanning sections".to_string(), - )); - } - - sections.push(desc); - if is_last { - break; - } - - offset = offset.saturating_add(advance); - } - - if sections.is_empty() { - return Err(Error::Invalid("no EWF sections found".to_string())); - } - - Ok(sections) -} - #[derive(Debug, Clone, Copy)] struct VolumeV1 { number_of_chunks: u32, @@ -1702,7 +1806,12 @@ fn parse_volume_like_section_v1( // For multi-segment EWF1, non-first segments may contain a `data` section that mirrors volume. let volume_desc = sections .iter() - .find(|s| s.type_string == "volume" || s.type_string == "disk" || s.type_string == "data") + .find(|s| { + matches!( + s.section_type, + Ewf1SectionType::Volume | Ewf1SectionType::Disk | Ewf1SectionType::Data + ) + }) .ok_or_else(|| { Error::Invalid("missing required section `volume` (or `disk`/`data`)".to_string()) })?; @@ -1839,13 +1948,13 @@ fn parse_chunk_groups_v1( let mut pending_sectors_end: Option = None; for desc in sections { - match desc.type_string.as_str() { + match &desc.section_type { // Chunk data section. The table that follows describes offsets into this region. - "sectors" | "sector" => { + Ewf1SectionType::Sectors | Ewf1SectionType::Sector => { let end = desc.start_offset.saturating_add(desc.size); pending_sectors_end = Some(end); } - x if x == table_type => { + _ if desc.section_type.as_str() == table_type => { let table = parse_table_section_v1(file, file_len, desc)?; if table.entries.is_empty() { return Err(Error::Invalid("table has no entries".to_string())); @@ -1933,7 +2042,7 @@ fn compute_chunk_data_end_offset_v1( ) -> Result { let last_chunk_data_offset = base_offset.saturating_add((last_entry & 0x7fff_ffff) as u64); - let end = if table_desc.type_string == "table2" { + let end = if matches!(table_desc.section_type, Ewf1SectionType::Table2) { // libewf: For table2 the chunk data is stored 2 sections before the table2 section. table_desc.start_offset.saturating_sub(table_desc.size) } else if last_chunk_data_offset < table_desc.start_offset { @@ -2125,34 +2234,6 @@ fn discover_segment_paths(base_path: &Path, naming: EwfxNaming) -> Result io::Result<()> { - let mut cur = offset; - while !buf.is_empty() { - #[cfg(unix)] - let n = file.read_at(buf, cur)?; - #[cfg(windows)] - let n = file.seek_read(buf, cur)?; - - if n == 0 { - return Err(io::Error::from(io::ErrorKind::UnexpectedEof)); - } - cur = cur.saturating_add(n as u64); - buf = &mut buf[n..]; - } - Ok(()) -} - -fn read_file_range(file: &File, file_len: u64, start: u64, end: u64) -> Result> { - if end > file_len || start >= end { - return Err(Error::Invalid("file range out of bounds".to_string())); - } - let len = usize::try_from(end - start) - .map_err(|_| Error::Invalid("range length overflow".to_string()))?; - let mut buf = vec![0u8; len]; - read_exact_at(file, start, &mut buf)?; - Ok(buf) -} - fn md5_file_range(file: &File, file_len: u64, start: u64, len: u64) -> Result<[u8; 16]> { if start >= file_len { return Err(Error::Invalid("file range out of bounds".to_string())); @@ -2180,13 +2261,6 @@ fn md5_file_range(file: &File, file_len: u64, start: u64, len: u64) -> Result<[u Ok(h.finalize().into()) } -// --- Parsing helpers --- - -fn parse_ascii_nul_terminated(bytes: &[u8]) -> String { - let len = bytes.iter().position(|&b| b == 0).unwrap_or(bytes.len()); - String::from_utf8_lossy(&bytes[..len]).to_string() -} - fn div_ceil_u64(a: u64, b: u64) -> u64 { if b == 0 { return 0; @@ -2194,20 +2268,6 @@ fn div_ceil_u64(a: u64, b: u64) -> u64 { a / b + u64::from(!a.is_multiple_of(b)) } -fn adler32_rfc1950(data: &[u8]) -> u32 { - // RFC1950 adler32; same as zlib's adler32. - const MOD_ADLER: u32 = 65521; - let mut a: u32 = 1; - let mut b: u32 = 0; - - for &byte in data { - a = (a + u32::from(byte)) % MOD_ADLER; - b = (b + a) % MOD_ADLER; - } - - (b << 16) | a -} - #[derive(Debug)] struct Ewf1SegmentParsed { path: PathBuf, @@ -2242,6 +2302,18 @@ pub struct LefEntry { pub size: u64, /// Data extents (empty for directories). pub extents: Vec, + /// File identifier (e.g. “inode” / metadata address), if present in the serialized tree. + /// + /// Note: some LEF encoders do not provide stable identifiers; in those cases this is `None`. + pub file_identifier: Option, + /// Access time (seconds since Unix epoch), if present in the serialized tree. + pub access_time: Option, + /// Modification time (seconds since Unix epoch), if present in the serialized tree. + pub modification_time: Option, + /// Entry modification (metadata change) time (seconds since Unix epoch), if present. + pub entry_modification_time: Option, + /// Creation time (seconds since Unix epoch), if present in the serialized tree. + pub creation_time: Option, } #[derive(Debug)] @@ -2313,7 +2385,7 @@ impl LefReader { let entry = self .entries() .iter() - .find(|e| e.path == want) + .find(|e| e.path == want && !e.is_dir) .ok_or_else(|| Error::Invalid(format!("file not found in LEF: `{path}`")))?; self.read_entry(entry) } @@ -2425,9 +2497,10 @@ fn open_l01(path: &Path) -> Result { ))); } - let sections = - parse_ewf1_section_descriptors(&file, file_len, hdr.sections_start_offset())?; - any_table2 |= sections.iter().any(|s| s.type_string == "table2"); + let sections = Ewf1SectionDescriptor::scan(&file, file_len, hdr.sections_start_offset())?; + any_table2 |= sections + .iter() + .any(|s| matches!(s.section_type, Ewf1SectionType::Table2)); parsed.push(Ewf1SegmentParsed { path: seg_path.clone(), @@ -2510,7 +2583,7 @@ fn open_lx01(path: &Path) -> Result { let section = last .sections .iter() - .find(|s| s.section_type == EWF2_SECTION_TYPE_SINGLE_FILES_DATA) + .find(|s| s.section_type == Ewf2SectionType::SingleFilesData) .ok_or_else(|| { Error::Invalid("missing EWF2 single files data section (0x20)".to_string()) })?; @@ -2560,7 +2633,7 @@ fn parse_chunk_geometry_v1_allow_zero_chunks( // bytes_per_sector still describe the chunk size. let candidate = ["data", "disk", "volume"] .into_iter() - .find_map(|name| sections.iter().find(|s| s.type_string == name)); + .find_map(|name| sections.iter().find(|s| s.section_type.as_str() == name)); let sec = candidate.ok_or_else(|| Error::Invalid("missing data/disk/volume section".to_string()))?; @@ -2593,7 +2666,7 @@ fn parse_ewf1_ltree( ) -> Result<(u64, Vec)> { let ltree = sections .iter() - .find(|s| s.type_string == "ltree") + .find(|s| matches!(s.section_type, Ewf1SectionType::LTree)) .ok_or_else(|| Error::Invalid("missing ltree section".to_string()))?; let (start, end) = ltree.data_range()?; if end.saturating_sub(start) < 48 { @@ -2822,7 +2895,47 @@ fn flatten_encase7_node(node: &EncaseNode, prefix: &str, out: &mut Vec format!("{prefix}/{name}") }; + let file_identifier = node + .values + .get("id") + .and_then(|s| s.trim().parse::().ok()); + + // EnCase “entry” timestamps appear as seconds since Unix epoch in the wild and in libewf’s + // test corpus (`ewf_test_single_files_data1`). Keep them as signed to allow sentinel values + // like `-1`. + let access_time = node + .values + .get("ac") + .and_then(|s| s.trim().parse::().ok()); + let modification_time = node + .values + .get("wr") + .and_then(|s| s.trim().parse::().ok()); + let entry_modification_time = node + .values + .get("mo") + .and_then(|s| s.trim().parse::().ok()); + let creation_time = node + .values + .get("cr") + .and_then(|s| s.trim().parse::().ok()); + if is_dir { + // Represent directories explicitly (except for the anonymous root) so callers can produce a + // full hierarchy listing and bodyfile output. + if !this_prefix.is_empty() { + out.push(LefEntry { + path: normalize_lef_path(&this_prefix), + is_dir: true, + size: 0, + extents: Vec::new(), + file_identifier, + access_time, + modification_time, + entry_modification_time, + creation_time, + }); + } for child in &node.children { flatten_encase7_node(child, &this_prefix, out); } @@ -2864,6 +2977,11 @@ fn flatten_encase7_node(node: &EncaseNode, prefix: &str, out: &mut Vec is_dir: false, size, extents, + file_identifier, + access_time, + modification_time, + entry_modification_time, + creation_time, }); } @@ -2890,9 +3008,21 @@ fn parse_binary_extents(value: &str) -> Result> { Ok(out) } +fn zlib_decompress_all(data: &[u8]) -> Result> { + let mut decoder = ZlibDecoder::new(data); + let mut out = Vec::new(); + decoder.read_to_end(&mut out)?; + Ok(out) +} + #[cfg(test)] mod tests { use super::*; + use crate::ewf1::EWF1_SECTION_DESCRIPTOR_SIZE; + use crate::ewf2::section::{ + EWF2_SECTION_DESCRIPTOR_SIZE, Ewf2SectionDataFlags, Ewf2SectionType, + make_ewf2_section_descriptor as make_ewf2_section_descriptor_bytes, + }; use std::io::Write as _; fn make_section_descriptor( @@ -2909,7 +3039,7 @@ mod tests { type_bytes[..copy_len].copy_from_slice(&src[..copy_len]); raw[..16].copy_from_slice(&type_bytes); - // next_offset (best-effort; not used by our scanner if size != 0) + // next_offset (informational; not used by our scanner if size != 0) let next_offset = start_offset.saturating_add(size); raw[16..24].copy_from_slice(&next_offset.to_le_bytes()); @@ -3006,42 +3136,23 @@ mod tests { pad as u32 } - fn make_ewf2_section_descriptor( - section_type: u32, - data_flags: u32, - previous_offset: u64, - data_size: u64, - padding_size: u32, - ) -> [u8; EWF2_SECTION_DESCRIPTOR_SIZE] { - let mut raw = [0u8; EWF2_SECTION_DESCRIPTOR_SIZE]; - raw[0..4].copy_from_slice(§ion_type.to_le_bytes()); - raw[4..8].copy_from_slice(&data_flags.to_le_bytes()); - raw[8..16].copy_from_slice(&previous_offset.to_le_bytes()); - raw[16..24].copy_from_slice(&data_size.to_le_bytes()); - raw[24..28].copy_from_slice(&(EWF2_SECTION_DESCRIPTOR_SIZE as u32).to_le_bytes()); - raw[28..32].copy_from_slice(&padding_size.to_le_bytes()); - // data_integrity_hash (16) + padding (12) left as zeros - let checksum = adler32_rfc1950(&raw[..EWF2_SECTION_DESCRIPTOR_SIZE - 4]); - raw[EWF2_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); - raw - } - fn append_ewf2_section_with_flags( file: &mut Vec, - section_type: u32, - data_flags: u32, + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, previous_offset: u64, data: Vec, padding_size_field: u32, ) -> u64 { file.extend_from_slice(&data); let desc_off = file.len() as u64; - let desc = make_ewf2_section_descriptor( + let desc = make_ewf2_section_descriptor_bytes( section_type, data_flags, previous_offset, data.len() as u64, padding_size_field, + [0u8; 16], ); file.extend_from_slice(&desc); desc_off @@ -3067,7 +3178,7 @@ mod tests { struct Ewf2SectorTableEntrySpec { offset: u64, size: u32, - flags: u32, + flags: Ewf2ChunkDataFlags, } fn build_ewf2_sector_table(first_chunk: u64, entries: &[Ewf2SectorTableEntrySpec]) -> Vec { @@ -3089,7 +3200,7 @@ mod tests { for e in entries { entries_bytes.extend_from_slice(&e.offset.to_le_bytes()); entries_bytes.extend_from_slice(&e.size.to_le_bytes()); - entries_bytes.extend_from_slice(&e.flags.to_le_bytes()); + entries_bytes.extend_from_slice(&e.flags.raw().to_le_bytes()); } table.extend_from_slice(&entries_bytes); @@ -3139,8 +3250,8 @@ mod tests { fn push_section_with_flags( &mut self, - section_type: u32, - data_flags: u32, + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, data: Vec, padding_size_field: u32, ) -> u64 { @@ -3158,17 +3269,22 @@ mod tests { fn push_section( &mut self, - section_type: u32, + section_type: Ewf2SectionType, data: Vec, padding_size_field: u32, ) -> u64 { - self.push_section_with_flags(section_type, 0, data, padding_size_field) + self.push_section_with_flags( + section_type, + Ewf2SectionDataFlags::new(0), + data, + padding_size_field, + ) } fn push_compressed_main_object( &mut self, - section_type: u32, - data_flags: u32, + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, pairs: &[(&str, u64)], ) -> u64 { let s = ewf2_main_object(pairs); @@ -3183,29 +3299,33 @@ mod tests { ("bp", u64::from(bytes_per_sector)), ("ts", number_of_sectors), ]; - self.push_compressed_main_object(EWF2_SECTION_TYPE_DEVICE_INFORMATION, 0, &pairs) + self.push_compressed_main_object( + Ewf2SectionType::DeviceInformation, + Ewf2SectionDataFlags::new(0), + &pairs, + ) } fn device_information_with_flags( &mut self, bytes_per_sector: u32, number_of_sectors: u64, - data_flags: u32, + data_flags: Ewf2SectionDataFlags, ) -> u64 { let pairs = [ ("bp", u64::from(bytes_per_sector)), ("ts", number_of_sectors), ]; - self.push_compressed_main_object( - EWF2_SECTION_TYPE_DEVICE_INFORMATION, - data_flags, - &pairs, - ) + self.push_compressed_main_object(Ewf2SectionType::DeviceInformation, data_flags, &pairs) } fn case_data(&mut self, sectors_per_chunk: u32, chunk_count: u64) -> u64 { let pairs = [("sb", u64::from(sectors_per_chunk)), ("tb", chunk_count)]; - self.push_compressed_main_object(EWF2_SECTION_TYPE_CASE_DATA, 0, &pairs) + self.push_compressed_main_object( + Ewf2SectionType::CaseData, + Ewf2SectionDataFlags::new(0), + &pairs, + ) } fn sector_data_uncompressed_chunk(&mut self, chunk: &[u8]) -> (u64, u32) { @@ -3225,17 +3345,21 @@ mod tests { let chunk_data_size = sector_data.len() as u32; let pad = pad16(&mut sector_data); - self.push_section(EWF2_SECTION_TYPE_SECTOR_DATA, sector_data, pad); + self.push_section(Ewf2SectionType::SectorData, sector_data, pad); (chunk_offset, chunk_data_size) } fn sector_table(&mut self, first_chunk: u64, entries: &[Ewf2SectorTableEntrySpec]) -> u64 { let table = build_ewf2_sector_table(first_chunk, entries); // libewf uses padding_size=24 for sector tables (12 after header + 12 after footer). - self.push_section(EWF2_SECTION_TYPE_SECTOR_TABLE, table, 24) + self.push_section(Ewf2SectionType::SectorTable, table, 24) } - fn utf16le_no_bom_section(&mut self, section_type: u32, text: &str) -> Ewf2WrittenSection { + fn utf16le_no_bom_section( + &mut self, + section_type: Ewf2SectionType, + text: &str, + ) -> Ewf2WrittenSection { let mut data = encode_utf16le_no_bom(text); let unpadded_len = data.len() as u64; let pad = pad16(&mut data); @@ -3248,11 +3372,11 @@ mod tests { } fn single_files_data(&mut self, ltree_text: &str) -> Ewf2WrittenSection { - self.utf16le_no_bom_section(EWF2_SECTION_TYPE_SINGLE_FILES_DATA, ltree_text) + self.utf16le_no_bom_section(Ewf2SectionType::SingleFilesData, ltree_text) } fn done(&mut self) -> u64 { - self.push_section(EWF2_SECTION_TYPE_DONE, Vec::new(), 0) + self.push_section(Ewf2SectionType::Done, Vec::new(), 0) } fn patch_descriptor_data_size(&mut self, desc_off: u64, data_size: u64) -> Result<()> { @@ -3397,12 +3521,12 @@ mod tests { &[Ewf2SectorTableEntrySpec { offset: chunk_data_offset, size: chunk_data_size, - flags: EWF2_CHUNK_DATA_FLAG_CHECKSUMED, + flags: Ewf2ChunkDataFlags::CHECKSUMED, }], ); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let img = EwfReader::open(&path)?; assert_eq!(img.len(), 512); @@ -3433,7 +3557,7 @@ mod tests { f.sector_table(0, &[]); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let err = EwfReader::open(&path).unwrap_err(); assert!(matches!(err, Error::Invalid(_))); @@ -3456,7 +3580,7 @@ mod tests { f.sector_table(0, &[]); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let err = EwfReader::open(&path).unwrap_err(); assert!(matches!(err, Error::Invalid(_))); @@ -3473,7 +3597,7 @@ mod tests { let chunk = vec![b'Z'; 512]; let mut f = Ewf2TestFile::new_ex01(set_id); - f.device_information_with_flags(512, 1, EWF2_SECTION_DATA_FLAG_ENCRYPTED); + f.device_information_with_flags(512, 1, Ewf2SectionDataFlags::ENCRYPTED); f.case_data(1, 1); let (chunk_data_offset, chunk_data_size) = f.sector_data_uncompressed_chunk(&chunk); f.sector_table( @@ -3481,12 +3605,12 @@ mod tests { &[Ewf2SectorTableEntrySpec { offset: chunk_data_offset, size: chunk_data_size, - flags: EWF2_CHUNK_DATA_FLAG_CHECKSUMED, + flags: Ewf2ChunkDataFlags::CHECKSUMED, }], ); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let err = EwfReader::open(&path).unwrap_err(); match err { @@ -3703,6 +3827,61 @@ mod tests { Ok(()) } + #[test] + fn test_parse_encase7_tree_extracts_entry_metadata_and_directories() -> Result<()> { + // This is a minimal EnCase 7-style tree with one directory ("dir") and one file + // ("dir/file.txt"), including basic metadata fields used by `ewfinfo` outputs. + // + // The field identifiers match the ones observed in libewf’s tests + // (`external/libewf/tests/ewf_test_single_files.c` / `ewf_test_ltree_section.c`): + // - id: file identifier + // - ac/wr/mo/cr: access/write/metadata-change/creation times (seconds since epoch) + let tree = concat!( + "2\n", + "rec\n", + "tb\n", + "5\n", + "\n", + "entry\n", + "1\t1\n", + "p\tn\tid\tac\twr\tmo\tcr\tls\tbe\n", + "0\t1\n", // root children + "1\n", // root values (p=1) + "0\t1\n", // dir children + "1\tdir\t1\t10\t20\t30\t40\t0\t\n", + "0\t0\n", // file children + "\tfile.txt\t42\t100\t200\t300\t400\t5\t1 0 5\n", + "\n", + ); + + let (_total, entries) = parse_encase7_tree(tree)?; + + let dir = entries.iter().find(|e| e.path == "dir").expect("dir"); + assert!(dir.is_dir); + assert_eq!(dir.size, 0); + assert!(dir.extents.is_empty()); + assert_eq!(dir.file_identifier, Some(1)); + assert_eq!(dir.access_time, Some(10)); + assert_eq!(dir.modification_time, Some(20)); + assert_eq!(dir.entry_modification_time, Some(30)); + assert_eq!(dir.creation_time, Some(40)); + + let file = entries + .iter() + .find(|e| e.path == "dir/file.txt") + .expect("dir/file.txt"); + assert!(!file.is_dir); + assert_eq!(file.size, 5); + assert_eq!(file.extents, vec![LefExtent { offset: 0, size: 5 }]); + assert_eq!(file.file_identifier, Some(42)); + assert_eq!(file.access_time, Some(100)); + assert_eq!(file.modification_time, Some(200)); + assert_eq!(file.entry_modification_time, Some(300)); + assert_eq!(file.creation_time, Some(400)); + + Ok(()) + } + #[test] fn test_lef_lx01_single_file_read() -> Result<()> { let chunk = hello_chunk_512(); @@ -3720,13 +3899,13 @@ mod tests { &[Ewf2SectorTableEntrySpec { offset: chunk_data_offset, size: chunk_data_size, - flags: EWF2_CHUNK_DATA_FLAG_CHECKSUMED, + flags: Ewf2ChunkDataFlags::CHECKSUMED, }], ); let _ = f.single_files_data(ENCASE7_TREE_SINGLE_HELLO_TXT); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let lef = LefReader::open(&path)?; let data = lef.read_file("hello.txt")?; @@ -3752,7 +3931,7 @@ mod tests { &[Ewf2SectorTableEntrySpec { offset: chunk_data_offset, size: chunk_data_size, - flags: EWF2_CHUNK_DATA_FLAG_CHECKSUMED, + flags: Ewf2ChunkDataFlags::CHECKSUMED, }], ); @@ -3762,7 +3941,7 @@ mod tests { // Add a "poison" section whose bytes look like a new (invalid) `rec` category. If we // over-read beyond the single-files section bounds, parsing will fail. let poison_text = "\n\nrec\nx\n"; - let poison = f.utf16le_no_bom_section(0xdead_beefu32, poison_text); + let poison = f.utf16le_no_bom_section(Ewf2SectionType::Unknown(0xdead_beef), poison_text); // Corrupt the single-files section descriptor's `data_size` to claim the section extends // past its descriptor and into the following section. @@ -3784,7 +3963,7 @@ mod tests { f.patch_descriptor_data_size(single.desc_off, inflated_data_size)?; f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let lef = LefReader::open(&path)?; let data = lef.read_file("hello.txt")?; @@ -3811,7 +3990,7 @@ mod tests { f.sector_table(0, &[]); let _ = f.single_files_data(ENCASE7_TREE_SINGLE_HELLO_TXT); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let err = LefReader::open(&path).unwrap_err(); assert!(matches!(err, Error::Invalid(_))); @@ -3835,10 +4014,145 @@ mod tests { let _ = f.single_files_data(ENCASE7_TREE_SINGLE_HELLO_TXT); f.done(); - std::fs::write(&path, &f.into_bytes())?; + std::fs::write(&path, f.into_bytes())?; let err = LefReader::open(&path).unwrap_err(); assert!(matches!(err, Error::Invalid(_))); Ok(()) } + + #[test] + fn test_parse_ewf1_header_ascii_av_maps_to_acquiry_software_version() -> Result<()> { + let header = concat!( + "1\r\n", + "main\r\n", + "c\tn\ta\te\tt\tav\tov\tm\tu\tp\tr\r\n", + "1\t1.1\tDesc\tExam\tNotes\tVERS\tLinux\t2006-01-02 03:04:05\t2006-01-02 03:04:05\t0\tn\r\n" + ); + + let mut out = crate::metadata::HeaderValues::default(); + crate::ewf1::header::Ewf1HeaderAscii::new(header.as_bytes()).parse_into(&mut out)?; + + assert_eq!(out.acquisition_software_version.as_deref(), Some("VERS")); + assert!(out.acquisition_software.is_none()); + Ok(()) + } + + #[test] + fn test_parse_ewf1_header2_av_maps_to_acquiry_software_version() -> Result<()> { + let header2 = concat!( + "1\n", + "main\n", + "a\tc\tn\te\tt\tmd\tsn\tav\tov\tm\tu\tp\n", + "Desc\t1\t1.1\tExam\tNotes\t\t\tVERS\tLinux\t2006-01-02 03:04:05\t2006-01-02 03:04:05\t0\n" + ); + + let mut out = crate::metadata::HeaderValues::default(); + let bytes = encode_utf16le_with_bom(header2); + crate::ewf1::header::Ewf1Header2Utf16Le::new(&bytes).parse_into(&mut out)?; + + assert_eq!(out.acquisition_software_version.as_deref(), Some("VERS")); + assert!(out.acquisition_software.is_none()); + Ok(()) + } + + #[test] + fn test_ewf1_image_metadata_extracts_header_volume_and_digests() -> Result<()> { + use crate::writer::{ + Ewf1CompressionLevel, Ewf1Format, EwfHeaderValues, EwfWriter, EwfWriterOptions, + }; + + let dir = tempfile::tempdir()?; + let path = dir.path().join("ewfinfo.E01"); + + let mut opts = EwfWriterOptions::new(Ewf1Format::E01, 1_474_560); + opts.bytes_per_sector = 512; + opts.sectors_per_chunk = 64; + opts.error_granularity = Some(7); + opts.compression_level = Ewf1CompressionLevel::None; + let set_id = [ + 0x86, 0x99, 0x10, 0xfc, 0xe1, 0x43, 0x49, 0x08, 0x93, 0x28, 0xaf, 0xed, 0xf4, 0xa7, + 0xbe, 0x1e, + ]; + opts.set_identifier = Some(set_id); + opts.header_values = EwfHeaderValues { + case_number: "1".to_string(), + evidence_number: "1.1".to_string(), + description: "Floppy".to_string(), + examiner_name: "John D.".to_string(), + notes: "Just a floppy in my system".to_string(), + acquisition_datetime: "2006-12-09 10:00:12".to_string(), + system_datetime: "2006-12-09 10:00:12".to_string(), + acquisition_software: "ewfacquire".to_string(), + acquisition_software_version: "20061209".to_string(), + acquisition_os: "Linux".to_string(), + }; + + let mut w = EwfWriter::create(&path, opts)?; + w.write(&vec![0u8; 1_474_560])?; + w.finish()?; + + let img = EwfReader::open(&path)?; + let meta = img.image_metadata()?; + + assert_eq!(meta.bytes_per_sector, 512); + assert_eq!(meta.sectors_per_chunk, 64); + assert_eq!(meta.error_granularity, 7); + assert_eq!(meta.header_values.case_number.as_deref(), Some("1")); + assert_eq!( + meta.header_values.acquisition_software.as_deref(), + Some("ewfacquire") + ); + assert_eq!( + meta.header_values.acquisition_software_version.as_deref(), + Some("20061209") + ); + + // Set identifier should be present. + assert_eq!(meta.set_identifier, Some(set_id)); + + // Digest hashes should be present. + assert!(meta.digests.md5.is_some()); + + Ok(()) + } + + #[test] + fn test_ewf2_image_metadata_extracts_set_id_geometry_and_hashes() -> Result<()> { + use crate::writer::{Ewf2CompressionMethod, Ewf2Writer, Ewf2WriterOptions}; + + let dir = tempfile::tempdir()?; + let path = dir.path().join("ewfinfo.Ex01"); + + let mut opts = Ewf2WriterOptions::new(32 * 1024); + opts.bytes_per_sector = 512; + opts.sectors_per_chunk = 64; + opts.error_granularity = Some(5); + opts.compression_method = Ewf2CompressionMethod::Zlib; + let set_id = [ + 0x11, 0x22, 0x33, 0x44, 0xaa, 0xbb, 0xcc, 0xdd, 0x00, 0x01, 0x02, 0x03, 0x10, 0x20, + 0x30, 0x40, + ]; + opts.set_identifier = Some(set_id); + + let mut w = Ewf2Writer::create(&path, opts)?; + w.write(&vec![0u8; 32 * 1024])?; + w.finish()?; + + let img = EwfReader::open(&path)?; + let meta = img.image_metadata()?; + + assert_eq!(meta.bytes_per_sector, 512); + assert_eq!(meta.sectors_per_chunk, 64); + assert_eq!(meta.error_granularity, 5); + + // Set identifier should be present. + assert_eq!(meta.set_identifier, Some(set_id)); + + // Hashes should be present. + assert!(meta.digests.md5.is_some()); + assert!(meta.digests.sha1.is_some()); + + Ok(()) + } } diff --git a/crates/ewf/src/util.rs b/crates/ewf/src/util.rs new file mode 100644 index 0000000..10aec16 --- /dev/null +++ b/crates/ewf/src/util.rs @@ -0,0 +1,65 @@ +//! Small, shared utilities used across EWF parsers and writers. +//! +//! This module intentionally stays lightweight and dependency-free. It exists to avoid duplicating +//! common low-level helpers (checksums, fixed-offset file reads, etc.) across the EWF1/EWF2 reader +//! and writer implementations. + +use std::fs::File; +use std::io; + +use crate::{Error, Result}; + +/// Reads exactly `buf.len()` bytes from `file` at a given absolute `offset`. +/// +/// This is a cross-platform wrapper around `FileExt` (pread/seek_read semantics). +pub(crate) fn read_exact_at(file: &File, offset: u64, mut buf: &mut [u8]) -> io::Result<()> { + #[cfg(unix)] + use std::os::unix::fs::FileExt as _; + #[cfg(windows)] + use std::os::windows::fs::FileExt as _; + + let mut cur = offset; + while !buf.is_empty() { + #[cfg(unix)] + let n = file.read_at(buf, cur)?; + #[cfg(windows)] + let n = file.seek_read(buf, cur)?; + + if n == 0 { + return Err(io::Error::from(io::ErrorKind::UnexpectedEof)); + } + cur = cur.saturating_add(n as u64); + buf = &mut buf[n..]; + } + Ok(()) +} + +/// Reads a file range into memory (half-open interval: `[start, end)`). +pub(crate) fn read_file_range(file: &File, file_len: u64, start: u64, end: u64) -> Result> { + if end > file_len || start >= end { + return Err(Error::Invalid("file range out of bounds".to_string())); + } + let len = usize::try_from(end - start) + .map_err(|_| Error::Invalid("range length overflow".to_string()))?; + let mut buf = vec![0u8; len]; + read_exact_at(file, start, &mut buf)?; + Ok(buf) +} + +/// Parses an ASCII NUL-terminated string from a fixed-width byte field. +pub(crate) fn parse_ascii_nul_terminated(bytes: &[u8]) -> String { + let len = bytes.iter().position(|&b| b == 0).unwrap_or(bytes.len()); + String::from_utf8_lossy(&bytes[..len]).to_string() +} + +/// Computes an Adler-32 checksum as defined by RFC1950 (zlib wrapper). +pub(crate) fn adler32_rfc1950(data: &[u8]) -> u32 { + const MOD_ADLER: u32 = 65521; + let mut a: u32 = 1; + let mut b: u32 = 0; + for &byte in data { + a = (a + u32::from(byte)) % MOD_ADLER; + b = (b + a) % MOD_ADLER; + } + (b << 16) | a +} diff --git a/crates/ewf/src/writer.rs b/crates/ewf/src/writer.rs index b073cef..00ea441 100644 --- a/crates/ewf/src/writer.rs +++ b/crates/ewf/src/writer.rs @@ -17,6 +17,21 @@ //! tooling (header2/header, volume/data, table/tables). More metadata surface area is added as part //! of the later EWF2/LEF/delta work. +use crate::ewf1::file_header::{Ewf1FileHeader, Ewf1Signature}; +use crate::ewf1::section::{ + Ewf1SectionDescriptor, Ewf1SectionType, make_ewf1_section_descriptor, make_ewf1_table_header, +}; +use crate::ewf1::volume as ewf1_volume; +use crate::ewf1::{ + EWF1_EVF_SIGNATURE, EWF1_FILE_HEADER_SIZE, EWF1_SECTION_DESCRIPTOR_SIZE, EWF1_TABLE_HEADER_SIZE, +}; +use crate::ewf2::chunk::Ewf2ChunkDataFlags; +use crate::ewf2::file_header::{EWF2_FILE_HEADER_SIZE, Ewf2FileHeader, Ewf2Kind}; +use crate::ewf2::section::{ + EWF2_SECTION_DESCRIPTOR_SIZE, Ewf2SectionDataFlags, Ewf2SectionType, + make_ewf2_section_descriptor, +}; +use crate::util::{adler32_rfc1950, read_exact_at, read_file_range}; use crate::{Error, Result}; use flate2::{Compression, write::ZlibEncoder}; use md5::{Digest as _, Md5}; @@ -26,39 +41,11 @@ use std::fs::File; use std::io::{self, Read as _, Seek, SeekFrom, Write as _}; use std::path::{Path, PathBuf}; -// EWF1 file header signature ("EVF\t\r\n\xff\0") -const EWF1_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; - -const EWF1_FILE_HEADER_SIZE: usize = 8 + 1 + 2 + 2; // 13 -const EWF1_SECTION_DESCRIPTOR_SIZE: usize = 16 + 8 + 8 + 40 + 4; // 76 -const EWF1_TABLE_HEADER_SIZE: usize = 4 + 4 + 8 + 4 + 4; // 24 - // --- EWF2 constants (EnCase 7 "EVF2"/".Ex01") --- -const EWF2_EVF_SIGNATURE: [u8; 8] = [0x45, 0x56, 0x46, 0x32, 0x0d, 0x0a, 0x81, 0x00]; // "EVF2\r\n\x81\0" - -const EWF2_FILE_HEADER_SIZE: usize = 32; -const EWF2_SECTION_DESCRIPTOR_SIZE: usize = 64; const EWF2_TABLE_HEADER_SIZE: usize = 32; // 20 bytes header + 12 bytes alignment padding const EWF2_TABLE_ENTRY_SIZE: usize = 16; const EWF2_TABLE_FOOTER_SIZE: usize = 16; // 4 bytes footer + 12 bytes alignment padding -const EWF2_SECTION_TYPE_DEVICE_INFORMATION: u32 = 0x0000_0001; -const EWF2_SECTION_TYPE_CASE_DATA: u32 = 0x0000_0002; -const EWF2_SECTION_TYPE_SECTOR_DATA: u32 = 0x0000_0003; -const EWF2_SECTION_TYPE_SECTOR_TABLE: u32 = 0x0000_0004; -const EWF2_SECTION_TYPE_MD5_HASH: u32 = 0x0000_0008; -const EWF2_SECTION_TYPE_SHA1_HASH: u32 = 0x0000_0009; -const EWF2_SECTION_TYPE_NEXT: u32 = 0x0000_000d; -const EWF2_SECTION_TYPE_DONE: u32 = 0x0000_000f; - -const EWF2_SECTION_DATA_FLAG_MD5HASHED: u32 = 0x0000_0001; -#[allow(dead_code)] -const EWF2_SECTION_DATA_FLAG_ENCRYPTED: u32 = 0x0000_0002; - -const EWF2_CHUNK_DATA_FLAG_COMPRESSED: u32 = 0x0000_0001; -const EWF2_CHUNK_DATA_FLAG_CHECKSUMED: u32 = 0x0000_0002; -const EWF2_CHUNK_DATA_FLAG_PATTERNFILL: u32 = 0x0000_0004; - /// EWF1 writer format (segment file naming + structural differences). #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum Ewf1Format { @@ -113,6 +100,170 @@ pub struct EwfHeaderValues { pub acquisition_os: String, } +impl EwfHeaderValues { + /// Build an EWF1 `header` section body (before zlib compression). + /// + /// This produces a minimal, stable EnCase-style header with CRLF line endings. + /// + /// References: + /// - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` + /// (“Header section”) + fn to_ewf1_header_ascii(&self, compression: Ewf1CompressionLevel) -> String { + // Minimal EnCase-like header structure: + // - 1 category (“main”) + // - identifiers line + values line + // Lines end with CRLF in EnCase-style headers. + // + // We keep the value set small but stable; tooling generally treats these as informational. + let mut s = String::new(); + s.push_str("1\r\n"); + s.push_str("main\r\n"); + s.push_str("c\tn\ta\te\tt\tav\tov\tm\tu\tp\tr\r\n"); + s.push_str(&format!( + "{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\r\n", + self.case_number, + self.evidence_number, + self.description, + self.examiner_name, + self.notes, + // `av` is the acquisition software version. + self.acquisition_software_version, + self.acquisition_os, + self.acquisition_datetime, + self.system_datetime, + "0", // password hash placeholder (no encryption for EWF1) + match compression { + Ewf1CompressionLevel::None => "n", + Ewf1CompressionLevel::Fast => "f", + Ewf1CompressionLevel::Best => "b", + } + )); + s.push_str("\r\n"); + s + } + + fn xml_escape(value: &str) -> String { + let mut out = String::with_capacity(value.len()); + for ch in value.chars() { + match ch { + '&' => out.push_str("&"), + '<' => out.push_str("<"), + '>' => out.push_str(">"), + '"' => out.push_str("""), + '\'' => out.push_str("'"), + _ => out.push(ch), + } + } + out + } + + /// Build an EWFX `xheader` section body (before zlib compression). + /// + /// The EWF spec documents `xheader` as UTF-8 XML. This is where libewf stores both + /// `acquiry_software` (name) and `acquiry_software_version`. + /// + /// References: + /// - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` + /// (“EWF-X” → “Xheader”) + fn to_ewfx_xheader_xml(&self) -> String { + let mut s = String::new(); + s.push_str("\n"); + s.push_str("\n"); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.case_number) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.description) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.examiner_name) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.evidence_number) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.notes) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.acquisition_os) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.acquisition_datetime) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.acquisition_software) + )); + s.push_str(&format!( + "\t{}\n", + Self::xml_escape(&self.acquisition_software_version) + )); + s.push_str("\n"); + s + } + + /// Build an EWF1 `header2` section body (before zlib compression). + /// + /// This produces a minimal EnCase 5–7 style header2: UTF-16LE with BOM and LF line endings. + /// + /// References: + /// - `external/libewf/documentation/Expert Witness Compression Format (EWF).asciidoc` + /// (“Header2 values”) + fn to_ewf1_header2_utf16le(&self) -> Vec { + // The full header2 semantics are extensive (categories, sources, subjects). We generate a + // structurally valid, minimal variant with empty categories beyond “main”. + let mut s = String::new(); + s.push_str("3\n"); + s.push_str("main\n"); + s.push_str("a\tc\tn\te\tt\tmd\tsn\tav\tov\tm\tu\tp\n"); + s.push_str(&format!( + "{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t0\n", + self.description, + self.case_number, + self.evidence_number, + self.examiner_name, + self.notes, + "", // media model + "", // serial + // `av` is the acquisition *software version*. + self.acquisition_software_version, + self.acquisition_os, + self.acquisition_datetime, + self.system_datetime, + )); + s.push('\n'); + + // srce category placeholder + s.push_str("srce\n"); + s.push_str("0 1\n"); + s.push_str("p\tn\tid\tev\ttb\tlo\tpo\tah\tsh\tgu\taq\n"); + s.push_str("0 0\n"); + s.push('\n'); + + // sub category placeholder + s.push_str("sub\n"); + s.push_str("0 1\n"); + s.push_str("p\tn\tid\tnu\tco\tgu\n"); + s.push_str("0 0\n"); + s.push('\n'); + + // UTF-16LE with BOM. + let mut out = Vec::with_capacity(2 + s.len() * 2); + out.extend_from_slice(&[0xff, 0xfe]); + for u in s.encode_utf16() { + out.extend_from_slice(&u.to_le_bytes()); + } + out + } +} + /// Options for creating or resuming an EWF1 writer. #[derive(Debug, Clone)] pub struct EwfWriterOptions { @@ -123,6 +274,10 @@ pub struct EwfWriterOptions { pub bytes_per_sector: u32, /// Sectors per chunk/block (typically 64, so chunk_size is 32768). pub sectors_per_chunk: u32, + /// The number of sectors to use as error granularity. + /// + /// If not set, this defaults to `sectors_per_chunk` (mirrors libewf’s acquisition tooling). + pub error_granularity: Option, /// Maximum size of a segment file in bytes (libewf default is 1500 MiB). pub segment_file_size: u64, /// Chunk compression level (E01 uses “compress if smaller”; S01 forces compression). @@ -142,6 +297,7 @@ impl EwfWriterOptions { media_size, bytes_per_sector: 512, sectors_per_chunk: 64, + error_granularity: None, segment_file_size: 1500 * 1024 * 1024, // libewf default compression_level: Ewf1CompressionLevel::default(), empty_block_compression: true, @@ -178,6 +334,10 @@ pub struct Ewf2WriterOptions { pub bytes_per_sector: u32, /// Sectors per chunk/block (typically 64, so chunk_size is 32768). pub sectors_per_chunk: u32, + /// The number of sectors to use as error granularity. + /// + /// If not set, this defaults to `sectors_per_chunk` (mirrors libewf’s acquisition tooling). + pub error_granularity: Option, /// Maximum size of a segment file in bytes. pub segment_file_size: u64, /// Segment-set compression method. @@ -196,6 +356,7 @@ impl Ewf2WriterOptions { media_size, bytes_per_sector: 512, sectors_per_chunk: 64, + error_granularity: None, segment_file_size: 1500 * 1024 * 1024, compression_method: Ewf2CompressionMethod::Zlib, pattern_fill: true, @@ -215,6 +376,7 @@ pub struct EwfWriter { // Media geometry (EWF1 volume/data sections). bytes_per_sector: u32, sectors_per_chunk: u32, + error_granularity: u32, number_of_sectors: u64, chunk_size: usize, chunk_count: u64, @@ -455,6 +617,10 @@ impl EwfWriter { let chunk_size = checked_chunk_size(sectors_per_chunk, bytes_per_sector)?; let number_of_sectors = checked_number_of_sectors(opts.media_size, bytes_per_sector)?; let chunk_count = div_ceil_u64(number_of_sectors, sectors_per_chunk as u64); + let mut error_granularity = opts.error_granularity.unwrap_or(sectors_per_chunk); + if error_granularity == 0 || error_granularity > sectors_per_chunk { + error_granularity = sectors_per_chunk; + } let base_path = remove_extension(path); let naming = Ewf1Naming::from_path(path, opts.format)?; @@ -473,6 +639,7 @@ impl EwfWriter { naming, bytes_per_sector, sectors_per_chunk, + error_granularity, number_of_sectors, chunk_size, chunk_count, @@ -668,43 +835,53 @@ impl EwfWriter { if segment_number == 0 || segment_number > u16::MAX as u32 { return Err(Error::Invalid("segment number out of bounds".to_string())); } - let file = self.file_mut()?; - file.write_all(&EWF1_EVF_SIGNATURE)?; - file.write_all(&[0x01])?; // start of fields - file.write_all(&(segment_number as u16).to_le_bytes())?; - file.write_all(&0u16.to_le_bytes())?; // end of fields + let hdr = Ewf1FileHeader::new(Ewf1Signature::Evf, segment_number as u16); + self.file_mut()?.write_all(&hdr.to_bytes())?; self.file_offset += EWF1_FILE_HEADER_SIZE as u64; Ok(()) } fn write_e01_header_sections(&mut self) -> Result<()> { // EnCase 4–7: header2 twice, then header once (all zlib-compressed). - let header2 = build_header2_utf16le(&self.opts.header_values); + let header2 = self.opts.header_values.to_ewf1_header2_utf16le(); let header2_z = zlib_compress(&header2, Compression::default())?; self.write_section_with_descriptor_v1("header2", &header2_z)?; self.write_section_with_descriptor_v1("header2", &header2_z)?; - let header = build_header_ascii(&self.opts.header_values, self.opts.compression_level); + let header = self + .opts + .header_values + .to_ewf1_header_ascii(self.opts.compression_level); let header_z = zlib_compress(header.as_bytes(), Compression::default())?; self.write_section_with_descriptor_v1("header", &header_z)?; + + // EWF-X (EWFX) XML header section. This is where libewf stores `acquiry_software` and + // `acquiry_software_version` as distinct values. + let xheader = self.opts.header_values.to_ewfx_xheader_xml(); + let xheader_z = zlib_compress(xheader.as_bytes(), Compression::default())?; + self.write_section_with_descriptor_v1("xheader", &xheader_z)?; Ok(()) } fn write_s01_header_section(&mut self) -> Result<()> { // SMART: a single header section, compressed using the same compression level as chunks. - let header = build_header_ascii(&self.opts.header_values, self.opts.compression_level); + let header = self + .opts + .header_values + .to_ewf1_header_ascii(self.opts.compression_level); let header_z = zlib_compress(header.as_bytes(), self.opts.compression_level.as_flate2())?; self.write_section_with_descriptor_v1("header", &header_z)?; Ok(()) } fn write_e01_volume_section(&mut self) -> Result<()> { - let data = build_volume_section_e01( + let data = ewf1_volume::build_volume_section_e01_1052( self.chunk_count, self.sectors_per_chunk, + self.error_granularity, self.bytes_per_sector, self.number_of_sectors, - self.opts.compression_level, + self.opts.compression_level.as_volume_byte(), self.set_identifier, ); self.write_section_with_descriptor_v1("volume", &data)?; @@ -712,12 +889,13 @@ impl EwfWriter { } fn write_e01_data_section(&mut self) -> Result<()> { - let data = build_volume_section_e01( + let data = ewf1_volume::build_volume_section_e01_1052( self.chunk_count, self.sectors_per_chunk, + self.error_granularity, self.bytes_per_sector, self.number_of_sectors, - self.opts.compression_level, + self.opts.compression_level.as_volume_byte(), self.set_identifier, ); self.write_section_with_descriptor_v1("data", &data)?; @@ -725,7 +903,7 @@ impl EwfWriter { } fn write_s01_volume_section(&mut self) -> Result<()> { - let data = build_volume_section_s01( + let data = ewf1_volume::build_volume_section_s01_94( self.chunk_count, self.sectors_per_chunk, self.bytes_per_sector, @@ -1081,8 +1259,8 @@ fn parse_segment_for_resume(path: &Path, format: Ewf1Format) -> Result "table2", @@ -1268,30 +1446,11 @@ fn make_section_descriptor_v1( next_offset: u64, size: u64, ) -> [u8; EWF1_SECTION_DESCRIPTOR_SIZE] { - let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; - - let mut type_bytes = [0u8; 16]; - let src = type_string.as_bytes(); - let copy_len = src.len().min(type_bytes.len().saturating_sub(1)); - type_bytes[..copy_len].copy_from_slice(&src[..copy_len]); - raw[..16].copy_from_slice(&type_bytes); - - raw[16..24].copy_from_slice(&next_offset.to_le_bytes()); - raw[24..32].copy_from_slice(&size.to_le_bytes()); - - // padding [32..72] left zero - let checksum = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); - raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); - raw + make_ewf1_section_descriptor(type_string, _start_offset, next_offset, size) } fn make_table_header_v1(number_of_entries: u32, base_offset: u64) -> [u8; EWF1_TABLE_HEADER_SIZE] { - let mut hdr = [0u8; EWF1_TABLE_HEADER_SIZE]; - hdr[0..4].copy_from_slice(&number_of_entries.to_le_bytes()); - hdr[8..16].copy_from_slice(&base_offset.to_le_bytes()); - let checksum = adler32_rfc1950(&hdr[..EWF1_TABLE_HEADER_SIZE - 4]); - hdr[EWF1_TABLE_HEADER_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); - hdr + make_ewf1_table_header(number_of_entries, base_offset) } fn build_table_section_v1(base_offset: u64, entries: &[u32]) -> Result> { @@ -1367,220 +1526,8 @@ fn build_hash_section(md5: &[u8; 16]) -> Vec { out } -fn build_volume_section_e01( - chunk_count: u64, - sectors_per_chunk: u32, - bytes_per_sector: u32, - number_of_sectors: u64, - compression_level: Ewf1CompressionLevel, - set_identifier: [u8; 16], -) -> Vec { - // FTK Imager / EnCase 1–7 / linen volume (1052 bytes) variant. - let mut out = vec![0u8; 1052]; - out[0] = 0x01; // fixed media - out[4..8].copy_from_slice(&(chunk_count as u32).to_le_bytes()); - out[8..12].copy_from_slice(§ors_per_chunk.to_le_bytes()); - out[12..16].copy_from_slice(&bytes_per_sector.to_le_bytes()); - out[16..24].copy_from_slice(&number_of_sectors.to_le_bytes()); - out[36] = 0x01; // media flags: “is an image file” - out[52] = compression_level.as_volume_byte(); - out[64..80].copy_from_slice(&set_identifier); - // checksum over [0..1048] - let checksum = adler32_rfc1950(&out[..1048]).to_le_bytes(); - out[1048..1052].copy_from_slice(&checksum); - out -} - -fn build_volume_section_s01( - chunk_count: u64, - sectors_per_chunk: u32, - bytes_per_sector: u32, - number_of_sectors: u64, -) -> Vec { - // EWF specification (94 bytes) variant used by SMART. - let mut out = vec![0u8; 94]; - out[0..4].copy_from_slice(&1u32.to_le_bytes()); // reserved (contains 0x01) - out[4..8].copy_from_slice(&(chunk_count as u32).to_le_bytes()); - out[8..12].copy_from_slice(§ors_per_chunk.to_le_bytes()); - out[12..16].copy_from_slice(&bytes_per_sector.to_le_bytes()); - out[16..20].copy_from_slice(&(number_of_sectors as u32).to_le_bytes()); - // signature at [85..90] is “SMART” in SMART files; we leave it zero to avoid guessing. - let checksum = adler32_rfc1950(&out[..90]).to_le_bytes(); - out[90..94].copy_from_slice(&checksum); - out -} - -fn build_header_ascii(values: &EwfHeaderValues, compression: Ewf1CompressionLevel) -> String { - // Minimal EnCase-like header structure: - // - 1 category (“main”) - // - identifiers line + values line - // Lines end with CRLF in EnCase-style headers. - // - // We keep the value set small but stable; tooling generally treats these as informational. - let mut s = String::new(); - s.push_str("1\r\n"); - s.push_str("main\r\n"); - s.push_str("c\tn\ta\te\tt\tav\tov\tm\tu\tp\tr\r\n"); - s.push_str(&format!( - "{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\r\n", - values.case_number, - values.evidence_number, - values.description, - values.examiner_name, - values.notes, - values.acquisition_software_version, - values.acquisition_os, - values.acquisition_datetime, - values.system_datetime, - "0", // password hash placeholder (no encryption for EWF1) - match compression { - Ewf1CompressionLevel::None => "n", - Ewf1CompressionLevel::Fast => "f", - Ewf1CompressionLevel::Best => "b", - } - )); - s.push_str("\r\n"); - s -} - -fn build_header2_utf16le(values: &EwfHeaderValues) -> Vec { - // Minimal EnCase 5–7 style header2: UTF-16LE text with BOM and LF line endings. - // - // The full header2 semantics are extensive (categories, sources, subjects). We generate a - // structurally valid, minimal variant with empty categories beyond “main”. - let mut s = String::new(); - s.push_str("3\n"); - s.push_str("main\n"); - s.push_str("a\tc\tn\te\tt\tmd\tsn\tav\tov\tm\tu\tp\n"); - s.push_str(&format!( - "{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t{}\t0\n", - values.description, - values.case_number, - values.evidence_number, - values.examiner_name, - values.notes, - "", // media model - "", // serial - values.acquisition_software, - values.acquisition_os, - values.acquisition_datetime, - values.system_datetime, - )); - s.push('\n'); - - // srce category placeholder - s.push_str("srce\n"); - s.push_str("0 1\n"); - s.push_str("p\tn\tid\tev\ttb\tlo\tpo\tah\tsh\tgu\taq\n"); - s.push_str("0 0\n"); - s.push('\n'); - - // sub category placeholder - s.push_str("sub\n"); - s.push_str("0 1\n"); - s.push_str("p\tn\tid\tnu\tco\tgu\n"); - s.push_str("0 0\n"); - s.push('\n'); - - // UTF-16LE with BOM. - let mut out = Vec::with_capacity(2 + s.len() * 2); - out.extend_from_slice(&[0xff, 0xfe]); - for u in s.encode_utf16() { - out.extend_from_slice(&u.to_le_bytes()); - } - out -} - -fn adler32_rfc1950(data: &[u8]) -> u32 { - const MOD_ADLER: u32 = 65521; - let mut a: u32 = 1; - let mut b: u32 = 0; - for &byte in data { - a = (a + u32::from(byte)) % MOD_ADLER; - b = (b + a) % MOD_ADLER; - } - (b << 16) | a -} - // --- Minimal parsing utilities reused by resume hashing --- -#[derive(Debug, Clone)] -struct Ewf1SectionDescriptor { - start_offset: u64, - type_string: String, - size: u64, -} - -impl Ewf1SectionDescriptor { - fn parse_at(file: &File, file_len: u64, start_offset: u64) -> Result { - if start_offset >= file_len { - return Err(io::Error::from(io::ErrorKind::UnexpectedEof).into()); - } - let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; - read_exact_at(file, start_offset, &mut raw)?; - - let stored = u32::from_le_bytes( - raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..] - .try_into() - .expect("len=4"), - ); - let calculated = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); - if stored != calculated { - return Err(Error::Corrupt( - "section descriptor checksum mismatch".to_string(), - )); - } - - let type_string = parse_ascii_nul_terminated(&raw[0..16]); - let next_offset = u64::from_le_bytes(raw[16..24].try_into().expect("len=8")); - let mut size = u64::from_le_bytes(raw[24..32].try_into().expect("len=8")); - if size == 0 && next_offset != start_offset && next_offset >= start_offset { - size = next_offset - start_offset; - } - Ok(Self { - start_offset, - type_string, - size, - }) - } - - fn data_range(&self) -> Result<(u64, u64)> { - let start = self.start_offset + EWF1_SECTION_DESCRIPTOR_SIZE as u64; - let end = self.start_offset + self.size; - Ok((start, end)) - } -} - -fn parse_ewf1_section_descriptors( - file: &File, - file_len: u64, - first_offset: u64, -) -> Result> { - let mut sections = Vec::new(); - let mut offset = first_offset; - for _ in 0..100_000 { - if offset == 0 || offset >= file_len { - break; - } - let desc = Ewf1SectionDescriptor::parse_at(file, file_len, offset)?; - let is_last = desc.type_string == "next" || desc.type_string == "done"; - let advance = if desc.size != 0 { - desc.size - } else { - EWF1_SECTION_DESCRIPTOR_SIZE as u64 - }; - sections.push(desc); - if is_last { - break; - } - offset = offset.saturating_add(advance); - } - if sections.is_empty() { - return Err(Error::Invalid("no EWF sections found".to_string())); - } - Ok(sections) -} - #[derive(Debug, Clone)] struct TableV1 { base_offset: u64, @@ -1651,11 +1598,11 @@ fn parse_chunk_groups_v1( let mut pending_sectors_end: Option = None; for desc in sections { - match desc.type_string.as_str() { - "sectors" | "sector" => { + match &desc.section_type { + Ewf1SectionType::Sectors | Ewf1SectionType::Sector => { pending_sectors_end = Some(desc.start_offset.saturating_add(desc.size)); } - x if x == table_type => { + _ if desc.section_type.as_str() == table_type => { let table = parse_table_section_v1(file, file_len, desc)?; let chunk_data_end = pending_sectors_end .take() @@ -1706,43 +1653,6 @@ fn chunk_range_v1(group: &Ewf1ChunkGroup, idx: usize) -> Result<(u64, u64, bool) Ok((start, end, is_compressed)) } -fn read_exact_at(file: &File, offset: u64, mut buf: &mut [u8]) -> io::Result<()> { - #[cfg(unix)] - use std::os::unix::fs::FileExt as _; - #[cfg(windows)] - use std::os::windows::fs::FileExt as _; - - let mut cur = offset; - while !buf.is_empty() { - #[cfg(unix)] - let n = file.read_at(buf, cur)?; - #[cfg(windows)] - let n = file.seek_read(buf, cur)?; - if n == 0 { - return Err(io::Error::from(io::ErrorKind::UnexpectedEof)); - } - cur += n as u64; - buf = &mut buf[n..]; - } - Ok(()) -} - -fn read_file_range(file: &File, file_len: u64, start: u64, end: u64) -> Result> { - if end > file_len || start >= end { - return Err(Error::Invalid("file range out of bounds".to_string())); - } - let len = - usize::try_from(end - start).map_err(|_| Error::Invalid("range overflow".to_string()))?; - let mut buf = vec![0u8; len]; - read_exact_at(file, start, &mut buf)?; - Ok(buf) -} - -fn parse_ascii_nul_terminated(bytes: &[u8]) -> String { - let len = bytes.iter().position(|&b| b == 0).unwrap_or(bytes.len()); - String::from_utf8_lossy(&bytes[..len]).to_string() -} - // === EWF2 writer (EWF2-Ex01 / EVF2) === #[derive(Debug, Clone, Copy)] @@ -1806,7 +1716,7 @@ impl Ewf2Naming { struct Ewf2TableEntry { offset_raw: [u8; 8], size: u32, - flags: u32, + flags: Ewf2ChunkDataFlags, } /// Streaming writer for EWF2-Ex01 images. @@ -1824,6 +1734,7 @@ pub struct Ewf2Writer { bytes_per_sector: u32, sectors_per_chunk: u32, + error_granularity: u32, number_of_sectors: u64, chunk_size: usize, chunk_count: u64, @@ -1877,6 +1788,10 @@ impl Ewf2Writer { let bytes_per_sector = opts.bytes_per_sector; let sectors_per_chunk = opts.sectors_per_chunk; let chunk_size = checked_chunk_size(sectors_per_chunk, bytes_per_sector)?; + let mut error_granularity = opts.error_granularity.unwrap_or(sectors_per_chunk); + if error_granularity == 0 || error_granularity > sectors_per_chunk { + error_granularity = sectors_per_chunk; + } if !opts.media_size.is_multiple_of(bytes_per_sector as u64) { return Err(Error::Invalid( @@ -1905,6 +1820,7 @@ impl Ewf2Writer { naming, bytes_per_sector, sectors_per_chunk, + error_granularity, number_of_sectors, chunk_size, chunk_count, @@ -2104,15 +2020,12 @@ impl Ewf2Writer { } fn write_ewf2_file_header(&mut self, segment_number: u32) -> Result<()> { - let compression_method = self.opts.compression_method.to_u16().to_le_bytes(); let set_id = self.set_identifier; + let compression_method = self.opts.compression_method.to_u16(); + let hdr = Ewf2FileHeader::new(Ewf2Kind::Ex01, compression_method, segment_number, set_id); let file = self.file_mut()?; - file.write_all(&EWF2_EVF_SIGNATURE)?; - file.write_all(&[2, 1])?; // major=2, minor=1 - file.write_all(&compression_method)?; - file.write_all(&segment_number.to_le_bytes())?; - file.write_all(&set_id)?; + file.write_all(&hdr.to_bytes())?; self.file_offset += EWF2_FILE_HEADER_SIZE as u64; Ok(()) } @@ -2135,8 +2048,8 @@ impl Ewf2Writer { }; let pad = pad16_bytes(&mut data); self.write_ewf2_section_with_descriptor( - EWF2_SECTION_TYPE_DEVICE_INFORMATION, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::DeviceInformation, + Ewf2SectionDataFlags::MD5_HASHED, pad, &data, )?; @@ -2147,6 +2060,7 @@ impl Ewf2Writer { let case_string = build_ewf2_case_data_string( self.chunk_count, self.sectors_per_chunk, + self.error_granularity, self.opts.compression_method.to_u16(), &self.opts.header_values, ); @@ -2162,8 +2076,8 @@ impl Ewf2Writer { }; let pad = pad16_bytes(&mut data); self.write_ewf2_section_with_descriptor( - EWF2_SECTION_TYPE_CASE_DATA, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::CaseData, + Ewf2SectionDataFlags::MD5_HASHED, pad, &data, )?; @@ -2176,26 +2090,26 @@ impl Ewf2Writer { self.table_entries.push(Ewf2TableEntry { offset_raw: [0u8; 8], // pattern = 0 size: 0, - flags: EWF2_CHUNK_DATA_FLAG_COMPRESSED | EWF2_CHUNK_DATA_FLAG_PATTERNFILL, + flags: Ewf2ChunkDataFlags::COMPRESSED | Ewf2ChunkDataFlags::PATTERNFILL, }); return Ok(()); } let data_offset = self.file_offset; - let (stored, flags): (Vec, u32) = match self.opts.compression_method { + let (stored, flags): (Vec, Ewf2ChunkDataFlags) = match self.opts.compression_method { Ewf2CompressionMethod::None => { // Uncompressed + Adler32 checksum. let mut v = Vec::with_capacity(self.chunk_size + 4); v.extend_from_slice(chunk); let checksum = adler32_rfc1950(chunk); v.extend_from_slice(&checksum.to_le_bytes()); - (v, EWF2_CHUNK_DATA_FLAG_CHECKSUMED) + (v, Ewf2ChunkDataFlags::CHECKSUMED) } Ewf2CompressionMethod::Zlib => { // Always compress; this matches libewf's behavior for formats that force compression. let z = zlib_compress_bytes(chunk)?; - (z, EWF2_CHUNK_DATA_FLAG_COMPRESSED) + (z, Ewf2ChunkDataFlags::COMPRESSED) } Ewf2CompressionMethod::Bzip2 => { return Err(Error::Unsupported( @@ -2244,8 +2158,8 @@ impl Ewf2Writer { let sector_md5: [u8; 16] = self.sector_data_md5.clone().finalize().into(); self.write_ewf2_section_descriptor_only( - EWF2_SECTION_TYPE_SECTOR_DATA, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::SectorData, + Ewf2SectionDataFlags::MD5_HASHED, self.sector_data_padding_total, sector_data_size, sector_md5, @@ -2257,8 +2171,8 @@ impl Ewf2Writer { &self.table_entries, )?; self.write_ewf2_section_with_descriptor( - EWF2_SECTION_TYPE_SECTOR_TABLE, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::SectorTable, + Ewf2SectionDataFlags::MD5_HASHED, 24, // libewf convention: 12 bytes after header + 12 bytes after footer &table_data, )?; @@ -2266,9 +2180,9 @@ impl Ewf2Writer { if last_segment { self.write_ewf2_hash_sections()?; // Done marker (no data). - self.write_ewf2_empty_section(EWF2_SECTION_TYPE_DONE)?; + self.write_ewf2_empty_section(Ewf2SectionType::Done)?; } else { - self.write_ewf2_empty_section(EWF2_SECTION_TYPE_NEXT)?; + self.write_ewf2_empty_section(Ewf2SectionType::Next)?; } // Close file. @@ -2282,16 +2196,16 @@ impl Ewf2Writer { let md5_data = build_ewf2_md5_section_data(&md5); self.write_ewf2_section_with_descriptor( - EWF2_SECTION_TYPE_MD5_HASH, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::Md5Hash, + Ewf2SectionDataFlags::MD5_HASHED, 12, &md5_data, )?; let sha1_data = build_ewf2_sha1_section_data(&sha1); self.write_ewf2_section_with_descriptor( - EWF2_SECTION_TYPE_SHA1_HASH, - EWF2_SECTION_DATA_FLAG_MD5HASHED, + Ewf2SectionType::Sha1Hash, + Ewf2SectionDataFlags::MD5_HASHED, 8, &sha1_data, )?; @@ -2299,19 +2213,25 @@ impl Ewf2Writer { Ok(()) } - fn write_ewf2_empty_section(&mut self, section_type: u32) -> Result<()> { + fn write_ewf2_empty_section(&mut self, section_type: Ewf2SectionType) -> Result<()> { let empty_md5: [u8; 16] = Md5::new().finalize().into(); - self.write_ewf2_section_descriptor_only(section_type, 0, 0, 0, empty_md5) + self.write_ewf2_section_descriptor_only( + section_type, + Ewf2SectionDataFlags::new(0), + 0, + 0, + empty_md5, + ) } fn write_ewf2_section_with_descriptor( &mut self, - section_type: u32, - data_flags: u32, + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, padding_size_field: u32, data: &[u8], ) -> Result<()> { - let md5_hash: [u8; 16] = if (data_flags & EWF2_SECTION_DATA_FLAG_MD5HASHED) != 0 { + let md5_hash: [u8; 16] = if data_flags.has_md5_integrity_hash() { let mut h = Md5::new(); h.update(data); h.finalize().into() @@ -2341,8 +2261,8 @@ impl Ewf2Writer { fn write_ewf2_section_descriptor_only( &mut self, - section_type: u32, - data_flags: u32, + section_type: Ewf2SectionType, + data_flags: Ewf2SectionDataFlags, padding_size_field: u32, data_size: u64, md5_hash: [u8; 16], @@ -2365,28 +2285,6 @@ impl Ewf2Writer { } } -fn make_ewf2_section_descriptor( - section_type: u32, - data_flags: u32, - previous_offset: u64, - data_size: u64, - padding_size: u32, - data_integrity_hash: [u8; 16], -) -> [u8; EWF2_SECTION_DESCRIPTOR_SIZE] { - let mut raw = [0u8; EWF2_SECTION_DESCRIPTOR_SIZE]; - raw[0..4].copy_from_slice(§ion_type.to_le_bytes()); - raw[4..8].copy_from_slice(&data_flags.to_le_bytes()); - raw[8..16].copy_from_slice(&previous_offset.to_le_bytes()); - raw[16..24].copy_from_slice(&data_size.to_le_bytes()); - raw[24..28].copy_from_slice(&(EWF2_SECTION_DESCRIPTOR_SIZE as u32).to_le_bytes()); - raw[28..32].copy_from_slice(&padding_size.to_le_bytes()); - raw[32..48].copy_from_slice(&data_integrity_hash); - // raw[48..60] padding is left as zeros. - let checksum = adler32_rfc1950(&raw[..EWF2_SECTION_DESCRIPTOR_SIZE - 4]); - raw[EWF2_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); - raw -} - fn encode_utf16le_with_bom(s: &str) -> Vec { let mut out = Vec::with_capacity(2 + s.len() * 2); out.extend_from_slice(&[0xff, 0xfe]); @@ -2421,11 +2319,17 @@ fn build_ewf2_device_information_string( fn build_ewf2_case_data_string( chunk_count: u64, sectors_per_chunk: u32, + error_granularity: u32, compression_method: u16, _values: &EwfHeaderValues, ) -> String { - // Minimal EWF2 case data: chunk count and chunk geometry, plus compression method. - format!("1\nmain\ntb\tsb\tcp\n{chunk_count}\t{sectors_per_chunk}\t{compression_method}\n") + // Minimal EWF2 case data: chunk count and chunk geometry, plus error granularity and compression method. + // + // libewf’s case-data object string includes many more tags; we keep a small subset required for + // `ewfinfo` and media geometry. + format!( + "1\nmain\ntb\tcp\tsb\tgr\n{chunk_count}\t{compression_method}\t{sectors_per_chunk}\t{error_granularity}\n" + ) } fn build_ewf2_sector_table_section_data( @@ -2456,7 +2360,7 @@ fn build_ewf2_sector_table_section_data( for e in entries { entries_bytes.extend_from_slice(&e.offset_raw); entries_bytes.extend_from_slice(&e.size.to_le_bytes()); - entries_bytes.extend_from_slice(&e.flags.to_le_bytes()); + entries_bytes.extend_from_slice(&e.flags.raw().to_le_bytes()); } out.extend_from_slice(&entries_bytes); diff --git a/crates/ewf/tests/golden/ewfinfo_image_dfxml.xml b/crates/ewf/tests/golden/ewfinfo_image_dfxml.xml new file mode 100644 index 0000000..9b0cc72 --- /dev/null +++ b/crates/ewf/tests/golden/ewfinfo_image_dfxml.xml @@ -0,0 +1,36 @@ + + + + Disk Image + Floppy + John D. + 1 + 1.1 + Just a floppy in my system + Sat Dec 9 10:00:12 2006 + + + ewfinfo + 0.1.0 + + rustc + + + + __OS__ + __ARCH__ + + + + __IMAGE_FILENAME__ + + + + + b37823c7a90d1917f719ba5927b23da8 + fe0786be0f9207ee934cc07917006ae7b8cb350a + + + 512 + + diff --git a/crates/ewf/tests/golden/ewfinfo_image_text.txt b/crates/ewf/tests/golden/ewfinfo_image_text.txt new file mode 100644 index 0000000..593a223 --- /dev/null +++ b/crates/ewf/tests/golden/ewfinfo_image_text.txt @@ -0,0 +1,36 @@ +Acquiry information +─────────────────── + Case number: 1 + Description: Floppy + Examiner name: John D. + Evidence number: 1.1 + Notes: Just a floppy in my system + Acquisition date: Sat Dec 9 10:00:12 2006 + System date: Sat Dec 9 10:00:12 2006 + Operating system used: Linux + Software used: ewfacquire + Software version used: 20061209 + Password: N/A + +EWF information +─────────────── + File format: EnCase 6 + Sectors per chunk: 64 + Error granularity: 1 + Compression method: deflate + Compression level: no compression + Set identifier: fc109986-43e1-0849-9328-afedf4a7be1e + +Media information +───────────────── + Media type: fixed disk + Is physical: no + Bytes per sector: 512 + Number of sectors: 2880 + Media size: 1.4 MiB (1474560 bytes) + +Digest hash information +─────────────────────── + MD5: b37823c7a90d1917f719ba5927b23da8 + SHA1: fe0786be0f9207ee934cc07917006ae7b8cb350a + diff --git a/crates/ewf/tests/golden/ewfinfo_logical_bodyfile.txt b/crates/ewf/tests/golden/ewfinfo_logical_bodyfile.txt new file mode 100644 index 0000000..6d181b7 --- /dev/null +++ b/crates/ewf/tests/golden/ewfinfo_logical_bodyfile.txt @@ -0,0 +1,2 @@ +0|dir|1|drwxrwxrwx|0|0|0|10.000000000|20.000000000|30.000000000|40.000000000 +0|dir/file.txt|42|-rwxrwxrwx|0|0|5|100.000000000|200.000000000|300.000000000|400.000000000 diff --git a/crates/ewf/tests/golden/ewfinfo_logical_file_entry.txt b/crates/ewf/tests/golden/ewfinfo_logical_file_entry.txt new file mode 100644 index 0000000..5a0919b --- /dev/null +++ b/crates/ewf/tests/golden/ewfinfo_logical_file_entry.txt @@ -0,0 +1,12 @@ +File entry information: + Name : dir/file.txt + Type : file + File identifier : 42 + Size : 5 + Access time : 100 + Modification time : 200 + Entry modification time : 300 + Creation time : 400 + +Extents: + 0 : offset=0 size=5 diff --git a/crates/ewf/tests/golden/ewfinfo_logical_hierarchy.txt b/crates/ewf/tests/golden/ewfinfo_logical_hierarchy.txt new file mode 100644 index 0000000..374a63c --- /dev/null +++ b/crates/ewf/tests/golden/ewfinfo_logical_hierarchy.txt @@ -0,0 +1,2 @@ +dir +dir/file.txt diff --git a/crates/ewf/tests/test_ewfinfo_golden.rs b/crates/ewf/tests/test_ewfinfo_golden.rs new file mode 100644 index 0000000..20070a3 --- /dev/null +++ b/crates/ewf/tests/test_ewfinfo_golden.rs @@ -0,0 +1,408 @@ +use std::process::Command; + +use ewf::writer::{Ewf1CompressionLevel, Ewf1Format, EwfHeaderValues, EwfWriter, EwfWriterOptions}; +use md5::{Digest as _, Md5}; + +const GOLDEN_IMAGE_TEXT: &str = include_str!("golden/ewfinfo_image_text.txt"); +const GOLDEN_IMAGE_DFXML: &str = include_str!("golden/ewfinfo_image_dfxml.xml"); +const GOLDEN_LOGICAL_HIERARCHY: &str = include_str!("golden/ewfinfo_logical_hierarchy.txt"); +const GOLDEN_LOGICAL_FILE_ENTRY: &str = include_str!("golden/ewfinfo_logical_file_entry.txt"); +const GOLDEN_LOGICAL_BODYFILE: &str = include_str!("golden/ewfinfo_logical_bodyfile.txt"); + +const EWF1_LVF_SIGNATURE: [u8; 8] = [0x4c, 0x56, 0x46, 0x09, 0x0d, 0x0a, 0xff, 0x00]; // "LVF\t\r\n\xff\0" +const EWF1_FILE_HEADER_SIZE: usize = 13; +const EWF1_SECTION_DESCRIPTOR_SIZE: usize = 76; +const EWF1_TABLE_HEADER_SIZE: usize = 24; + +fn normalize_newlines(s: &str) -> String { + s.replace("\r\n", "\n") +} + +fn ensure_trailing_newline(mut s: String) -> String { + if !s.ends_with('\n') { + s.push('\n'); + } + s +} + +fn strip_text_header(s: &str) -> &str { + // libewf (and our CLI) emits a text header: + // ewfinfo \n\n + if s.starts_with("ewfinfo ") + && let Some(idx) = s.find("\n\n") + { + return &s[idx + 2..]; + } + s +} + +fn normalize_text_output(stdout: &str) -> String { + let s = normalize_newlines(stdout); + ensure_trailing_newline(strip_text_header(&s).to_string()) +} + +fn replace_all_tag_text(s: &str, tag: &str, replacement: &str) -> String { + let open = format!("<{tag}>"); + let close = format!(""); + + let mut out = String::with_capacity(s.len()); + let mut rest = s; + + while let Some(open_idx) = rest.find(&open) { + out.push_str(&rest[..open_idx]); + out.push_str(&open); + + let after_open = &rest[open_idx + open.len()..]; + let Some(close_idx) = after_open.find(&close) else { + // Malformed input; keep the remainder unchanged. + out.push_str(after_open); + return out; + }; + + out.push_str(replacement); + out.push_str(&close); + rest = &after_open[close_idx + close.len()..]; + } + + out.push_str(rest); + out +} + +fn normalize_dfxml(stdout: &str) -> String { + let s = normalize_newlines(stdout); + let s = replace_all_tag_text(&s, "os_sysname", "__OS__"); + let s = replace_all_tag_text(&s, "arch", "__ARCH__"); + replace_all_tag_text(&s, "image_filename", "__IMAGE_FILENAME__") +} + +fn run_ewfinfo(args: &[&str]) -> std::process::Output { + let exe = env!("CARGO_BIN_EXE_ewfinfo"); + Command::new(exe).args(args).output().expect("run ewfinfo") +} + +fn build_synthetic_e01(dir: &tempfile::TempDir) -> std::path::PathBuf { + let path = dir.path().join("ewfinfo.E01"); + + let mut opts = EwfWriterOptions::new(Ewf1Format::E01, 1_474_560); + opts.bytes_per_sector = 512; + opts.sectors_per_chunk = 64; + // Make error granularity intentionally differ from sectors_per_chunk to catch copy/paste bugs + // in `ewfinfo` rendering. + opts.error_granularity = Some(1); + opts.compression_level = Ewf1CompressionLevel::None; + opts.set_identifier = Some([ + 0x86, 0x99, 0x10, 0xfc, 0xe1, 0x43, 0x49, 0x08, 0x93, 0x28, 0xaf, 0xed, 0xf4, 0xa7, 0xbe, + 0x1e, + ]); + opts.header_values = EwfHeaderValues { + case_number: "1".to_string(), + evidence_number: "1.1".to_string(), + description: "Floppy".to_string(), + examiner_name: "John D.".to_string(), + notes: "Just a floppy in my system".to_string(), + acquisition_datetime: "2006-12-09 10:00:12".to_string(), + system_datetime: "2006-12-09 10:00:12".to_string(), + acquisition_software: "ewfacquire".to_string(), + acquisition_software_version: "20061209".to_string(), + acquisition_os: "Linux".to_string(), + }; + + let mut w = EwfWriter::create(&path, opts).expect("create E01"); + w.write(&vec![0u8; 1_474_560]).expect("write"); + w.finish().expect("finish"); + path +} + +fn adler32_rfc1950(data: &[u8]) -> u32 { + const MOD_ADLER: u32 = 65521; + let mut a: u32 = 1; + let mut b: u32 = 0; + + for &byte in data { + a = (a + u32::from(byte)) % MOD_ADLER; + b = (b + a) % MOD_ADLER; + } + + (b << 16) | a +} + +fn make_section_descriptor( + type_string: &str, + start_offset: u64, + size: u64, +) -> [u8; EWF1_SECTION_DESCRIPTOR_SIZE] { + let mut raw = [0u8; EWF1_SECTION_DESCRIPTOR_SIZE]; + + // type string (ASCII, NUL-terminated) + let mut type_bytes = [0u8; 16]; + let src = type_string.as_bytes(); + let copy_len = src.len().min(type_bytes.len().saturating_sub(1)); + type_bytes[..copy_len].copy_from_slice(&src[..copy_len]); + raw[..16].copy_from_slice(&type_bytes); + + // next_offset (informational) + let next_offset = start_offset.saturating_add(size); + raw[16..24].copy_from_slice(&next_offset.to_le_bytes()); + + // size + raw[24..32].copy_from_slice(&size.to_le_bytes()); + + let checksum = adler32_rfc1950(&raw[..EWF1_SECTION_DESCRIPTOR_SIZE - 4]); + raw[EWF1_SECTION_DESCRIPTOR_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); + raw +} + +fn make_table_header(number_of_entries: u32, base_offset: u64) -> [u8; EWF1_TABLE_HEADER_SIZE] { + let mut hdr = [0u8; EWF1_TABLE_HEADER_SIZE]; + hdr[0..4].copy_from_slice(&number_of_entries.to_le_bytes()); + hdr[8..16].copy_from_slice(&base_offset.to_le_bytes()); + let checksum = adler32_rfc1950(&hdr[..EWF1_TABLE_HEADER_SIZE - 4]); + hdr[EWF1_TABLE_HEADER_SIZE - 4..].copy_from_slice(&checksum.to_le_bytes()); + hdr +} + +fn write_lvf_header(file: &mut Vec, segment_number: u16) { + file.extend_from_slice(&EWF1_LVF_SIGNATURE); + file.push(0x01); // start of fields + file.extend_from_slice(&segment_number.to_le_bytes()); + file.extend_from_slice(&0u16.to_le_bytes()); // end of fields + assert_eq!(file.len(), EWF1_FILE_HEADER_SIZE); +} + +fn encode_utf16le_no_bom(s: &str) -> Vec { + let mut out = Vec::with_capacity(s.len() * 2); + for u in s.encode_utf16() { + out.extend_from_slice(&u.to_le_bytes()); + } + out +} + +fn hello_chunk_512() -> [u8; 512] { + let mut out = [0u8; 512]; + out[..5].copy_from_slice(b"hello"); + out +} + +fn build_synthetic_l01(dir: &tempfile::TempDir) -> std::path::PathBuf { + let chunk = hello_chunk_512(); + + let mut file: Vec = Vec::new(); + write_lvf_header(&mut file, 1); + + let mut append_section = |typ: &str, body: &[u8]| -> u64 { + let start_offset = file.len() as u64; + let size = (EWF1_SECTION_DESCRIPTOR_SIZE + body.len()) as u64; + let desc = make_section_descriptor(typ, start_offset, size); + file.extend_from_slice(&desc); + file.extend_from_slice(body); + start_offset + }; + + // data section: number_of_chunks is 0 for L01, but chunk geometry is still present. + let mut data_body = vec![0u8; 24]; + data_body[0..4].copy_from_slice(&1u32.to_le_bytes()); // version/unknown + data_body[4..8].copy_from_slice(&0u32.to_le_bytes()); // number_of_chunks (often 0) + data_body[8..12].copy_from_slice(&1u32.to_le_bytes()); // sectors_per_chunk + data_body[12..16].copy_from_slice(&512u32.to_le_bytes()); // bytes_per_sector + data_body[16..24].copy_from_slice(&1u64.to_le_bytes()); // number_of_sectors + append_section("data", &data_body); + + // sectors: uncompressed chunk + Adler32 of chunk bytes + let mut sectors_body = Vec::new(); + sectors_body.extend_from_slice(&chunk); + let checksum = adler32_rfc1950(&chunk); + sectors_body.extend_from_slice(&checksum.to_le_bytes()); + let sectors_start = append_section("sectors", §ors_body); + let chunk_file_off = (sectors_start + EWF1_SECTION_DESCRIPTOR_SIZE as u64) as u32; + + // table2: one entry, base_offset=0, no compression flag. + let mut table2_body: Vec = Vec::new(); + table2_body.extend_from_slice(&make_table_header(1, 0)); + table2_body.extend_from_slice(&chunk_file_off.to_le_bytes()); + append_section("table2", &table2_body); + + // ltree: EnCase 7 style serialized tree (UTF-16LE without BOM). + let ltree_text = concat!( + "2\n", + "rec\n", + "tb\n", + "5\n", + "\n", + "entry\n", + "1\t1\n", + "p\tn\tid\tac\twr\tmo\tcr\tls\tbe\n", + "0\t1\n", + "1\n", + "0\t1\n", + "1\tdir\t1\t10\t20\t30\t40\t0\t\n", + "0\t0\n", + "\tfile.txt\t42\t100\t200\t300\t400\t5\t1 0 5\n", + "\n", + ); + let ltree_data = encode_utf16le_no_bom(ltree_text); + + let mut ltree_hdr = [0u8; 48]; + let md5 = { + let mut h = Md5::new(); + h.update(<ree_data); + let d = h.finalize(); + let mut out = [0u8; 16]; + out.copy_from_slice(&d[..]); + out + }; + ltree_hdr[0..16].copy_from_slice(&md5); + ltree_hdr[16..24].copy_from_slice(&(ltree_data.len() as u64).to_le_bytes()); + // checksum at 24..28 filled later + let mut hdr_for_checksum = ltree_hdr; + hdr_for_checksum[24..28].fill(0); + let hdr_checksum = adler32_rfc1950(&hdr_for_checksum); + ltree_hdr[24..28].copy_from_slice(&hdr_checksum.to_le_bytes()); + + let mut ltree_body = Vec::new(); + ltree_body.extend_from_slice(<ree_hdr); + ltree_body.extend_from_slice(<ree_data); + append_section("ltree", <ree_body); + + append_section("done", &[]); + + let path = dir.path().join("case.L01"); + std::fs::write(&path, &file).expect("write L01"); + path +} + +#[test] +fn test_ewfinfo_real_fixture_nps_formats_epoch_dates_and_reports_encase6() { + // Regression test for real-world EWF1 images where header dates are stored as Unix epoch + // seconds. libewf formats these according to `-d` (default: `ctime`) and reports `.E01` as + // "EnCase 6". + // + // We run the binary with a fixed TZ to keep output deterministic across CI environments. + let exe = env!("CARGO_BIN_EXE_ewfinfo"); + let fixture = std::path::PathBuf::from(env!("CARGO_MANIFEST_DIR")) + .join("../../testdata/ewf/nps-2010-emails.E01"); + assert!(fixture.exists(), "missing fixture: {}", fixture.display()); + + let out = Command::new(exe) + .env("TZ", "UTC") + .arg(&fixture) + .output() + .expect("run ewfinfo"); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let actual = normalize_text_output(&String::from_utf8_lossy(&out.stdout)); + + // Ensure we did not print the raw epoch value. + assert!(!actual.contains("1296677487")); + // Ensure the formatted ctime string is present (TZ=UTC). + assert!(actual.contains("Acquisition date: Wed Feb 2 20:11:27 2011")); + assert!(actual.contains("System date: Wed Feb 2 20:11:27 2011")); + // Ensure file format matches libewf. + assert!(actual.contains("File format: EnCase 6")); +} + +#[test] +fn test_ewfinfo_image_text_matches_golden() { + let dir = tempfile::tempdir().unwrap(); + let e01 = build_synthetic_e01(&dir); + + let out = run_ewfinfo(&[e01.to_str().unwrap()]); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let actual = normalize_text_output(&String::from_utf8_lossy(&out.stdout)); + let expected = ensure_trailing_newline(normalize_newlines(GOLDEN_IMAGE_TEXT)); + assert_eq!(actual, expected); +} + +#[test] +fn test_ewfinfo_image_dfxml_matches_golden() { + let dir = tempfile::tempdir().unwrap(); + let e01 = build_synthetic_e01(&dir); + + let out = run_ewfinfo(&["-f", "dfxml", e01.to_str().unwrap()]); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let actual = normalize_dfxml(&String::from_utf8_lossy(&out.stdout)); + let expected = normalize_newlines(GOLDEN_IMAGE_DFXML); + assert_eq!(actual, expected); +} + +#[test] +fn test_ewfinfo_logical_hierarchy_matches_golden() { + let dir = tempfile::tempdir().unwrap(); + let l01 = build_synthetic_l01(&dir); + + let out = run_ewfinfo(&["-H", l01.to_str().unwrap()]); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let actual = normalize_text_output(&String::from_utf8_lossy(&out.stdout)); + let expected = ensure_trailing_newline(normalize_newlines(GOLDEN_LOGICAL_HIERARCHY)); + assert_eq!(actual, expected); +} + +#[test] +fn test_ewfinfo_logical_file_entry_matches_golden() { + let dir = tempfile::tempdir().unwrap(); + let l01 = build_synthetic_l01(&dir); + + let out = run_ewfinfo(&["-F", "dir/file.txt", l01.to_str().unwrap()]); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let actual = normalize_text_output(&String::from_utf8_lossy(&out.stdout)); + let expected = ensure_trailing_newline(normalize_newlines(GOLDEN_LOGICAL_FILE_ENTRY)); + assert_eq!(actual, expected); +} + +#[test] +fn test_ewfinfo_logical_bodyfile_matches_golden() { + let dir = tempfile::tempdir().unwrap(); + let l01 = build_synthetic_l01(&dir); + + let bodyfile_path = dir.path().join("bodyfile"); + let out = run_ewfinfo(&[ + "-B", + bodyfile_path.to_str().unwrap(), + "-H", + l01.to_str().unwrap(), + ]); + assert!( + out.status.success(), + "stderr: {}", + String::from_utf8_lossy(&out.stderr) + ); + + let bodyfile = std::fs::read_to_string(&bodyfile_path).expect("read bodyfile"); + let actual = ensure_trailing_newline(normalize_newlines(&bodyfile)); + let expected = ensure_trailing_newline(normalize_newlines(GOLDEN_LOGICAL_BODYFILE)); + assert_eq!(actual, expected); +} + +#[test] +fn test_ewfinfo_logical_dfxml_is_not_implemented_yet() { + let dir = tempfile::tempdir().unwrap(); + let l01 = build_synthetic_l01(&dir); + + let out = run_ewfinfo(&["-f", "dfxml", "-H", l01.to_str().unwrap()]); + assert!(!out.status.success()); + let stderr = String::from_utf8_lossy(&out.stderr); + assert!(stderr.contains("dfxml output is not yet implemented")); +} diff --git a/external/refs/README.md b/external/refs/README.md index 93375dd..fd1a184 100644 --- a/external/refs/README.md +++ b/external/refs/README.md @@ -13,4 +13,12 @@ is compiled or linked into the Rust crates. This crate does **not** vendor the full upstream repository snapshot; the pinned commit file is enough to reproduce the reference checkout locally when needed. +### `dfxml-working-group/dfxml_schema` + +- **Pinned commit**: `external/refs/repos/dfxml-working-group__dfxml_schema.commit` +- **Upstream**: `https://github.com/dfxml-working-group/dfxml_schema` + +This repository is used as the reference for DFXML schema compliance. The workspace vendors a +minimal set of schema files under `crates/dfxml/schema/` for offline validation in tests. + diff --git a/external/refs/repos/dfxml-working-group__dfxml_schema.commit b/external/refs/repos/dfxml-working-group__dfxml_schema.commit new file mode 100644 index 0000000..d1f2e3b --- /dev/null +++ b/external/refs/repos/dfxml-working-group__dfxml_schema.commit @@ -0,0 +1,2 @@ +a80a8619d83ac23c7cd41bd9262be86fcebe3bd8 + diff --git a/testdata/ewf/nps-2010-emails.E01 b/testdata/ewf/nps-2010-emails.E01 new file mode 100644 index 0000000..6c87fa4 --- /dev/null +++ b/testdata/ewf/nps-2010-emails.E01 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c9ffd969954c2f9b9f97f459916c3d2e8755f596eda952c306ab3f9bc0d43bf1 +size 518680