Files
actual/.gitattributes
4e263d983d [AI] Detect CSV file encoding and add a per-account encoding selector (#8924)
* [AI] Detect CSV file encoding and add a per-account encoding selector

Read CSV/TSV files as raw bytes and decode them based on the detected
encoding instead of always assuming UTF-8: the byte order mark, BOM-less
UTF-16 (NUL byte analysis), strict UTF-8 validation, and a windows-1252
fallback that also decodes ISO-8859-1 content. A mostly-valid UTF-8 file
with a few corrupted bytes still decodes as UTF-8 with replacement
characters instead of falling back to windows-1252.

Add an encoding selector to the CSV import options, persisted per
account like the delimiter, for files in encodings that cannot be
detected automatically (e.g. windows-1250, ISO-8859-2).

Adds fixtures and tests for UTF-16 (with and without BOM), UTF-8 with
BOM, corrupted UTF-8, NUL-padded UTF-8, and windows-1252 content.

Fixes #6327

* Update VRT screenshots

Auto-generated by VRT workflow

PR: #8924

* Update VRT screenshots

Auto-generated by VRT workflow

PR: #8924

* [AI] Address review: require dominant NUL parity and translate encoding labels

Require one parity to clearly dominate (>= 90% of NUL bytes) before
detecting BOM-less UTF-16, so a UTF-8 file with dense contiguous NUL
padding (even split across parities) is no longer misdetected as
UTF-16. Adds a fixture with padding above the 10% density threshold.

Wrap the encoding selector labels in t() per the repository's
translated user-facing text requirement.

* [AI] Drop iso-8859-1 from the CSV encoding selector

TextDecoder resolves the iso-8859-1 label to the windows-1252 decoder
per the WHATWG encoding spec, so the option was redundant and could
mislead users into expecting true ISO-8859-1 decoding. windows-1252
already covers ISO-8859-1 content; keep the alias accepted in
decodeCsvBytes with a clarifying comment and an alias test.

* [AI] Make the iso-8859-1 alias test distinguish Windows-1252

The alias test now uses a fixture containing the 0x80 byte, which
windows-1252 decodes as the euro sign while a true ISO-8859-1 decoder
would produce a C1 control character, so the assertion actually
verifies the windows-1252 alias semantics.

* [AI] Simplify CSV encoding detection to BOM and explicit selection

Follow the review suggestion: the byte order mark selects UTF-16
LE/BE and everything else decodes as UTF-8; other encodings are
handled by the manual per-account selector this PR adds. Drop the
speculative BOM-less UTF-16, windows-1252 fallback, and corrupted
UTF-8 heuristics along with their fixtures and tests.

---------

Co-authored-by: François Lafleur <[email protected]>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 22:31:04 +00:00

26 lines
684 B
Plaintext

# Set the default behavior, in case people don't have core.autocrlf set.
* text=auto
# Explicitly declare text files you want to always be normalized and converted
# to native line endings on checkout.
# *.c text
# *.h text
# Declare files that will always have LF line endings on checkout.
*.js text eol=lf
*.ts text eol=lf
*.sh text eol=lf
*.tsx text eol=lf
**/bin/* text eol=lf
yarn.lock text eol=lf
# Declare files that will always have CRLF line endings on checkout.
packages/mobile-client/android/gradlew.bat text eol=crlf
# Denote all files that are truly binary and should not be modified.
*.png binary
*.jpg binary
packages/loot-core/src/mocks/files/utf-16*.csv binary