Speed up project search over misclassified binary files (#64549)

Based on
https://github.com/zed-industries/zed/issues/38799#issuecomment-5758593194
finding.

Project search treated any NUL-free file as text, then read it whole,
ran chardetng over every byte (~34 MiB/s) and decoded it into a String
before searching. For raw assets (.yuv/.raw/.onnx/.glb/.tif in #38799)
that cost ~8 s and ~2.6× the file size in RAM per worker.

- Bound chardetng to a 1 MiB window starting at the first non-ASCII
byte.
- Stream-decode legacy/UTF-16 files in project search (`DecodingReader`)
instead of materializing them.
- Scan the single-line text pre-filter in 64 KiB blocks instead of
`read_line`.
- Run the control-byte check even without NULs; treat all high bytes as
text so legacy encodings (windows-1251, GBK, ...) stay text.
- Add magic numbers for TIFF, glTF, ELF, Mach-O, gzip/xz/zstd, SQLite,
HDF5, NumPy, safetensors, GGUF, fonts, media containers and others.
- Move the classifier into a `file_content` crate and use it for git
commit blobs too, replacing the separate NUL-only `is_binary_content`.

Measured on a 200 MiB NUL-free raw file: 8.5 s → ~1.5 s per file, no
whole-file allocation.

Release Notes:

- Improved project search speed and memory usage over misclassified
binary files
8bbcd83968Kirill Bulatov committed on 9/30/2026, 5:39:08 PM· committed by GitHubparentd284cbb
20 files changedLine totals unavailable