katlab tools/extract/zip-filename-encoding support on Ko-fi

Fix Garbled File Names in .zip Files

Japanese, Chinese, Korean or Cyrillic file names that come out as gibberish are a known ZIP problem. Open the archive here, check the detected encoding, and save a clean copy with UTF-8 names.

Files are processed locally. Nothing is uploaded to a server.
Shift_JISGBK / GB18030Big5EUC-KRCP866Windows-1251KOI8-R

How to fix garbled file names in a ZIP

  1. Drop the .zip file onto the box above. The tool checks the stored names and picks the most likely encoding.
  2. Look at the file list. If the names still look wrong, choose another entry in the Filename encoding dropdown; the list updates as a live preview.
  3. When the names read correctly, download all files as a .zip. The new archive stores the names in UTF-8.

Why ZIP file names turn into gibberish

The ZIP format is older than Unicode on the desktop. For most of its history, a ZIP tool wrote file names as raw bytes in whatever code page the computer was set to: Shift_JIS on a Japanese system, GBK on a Simplified Chinese one, Big5 for Traditional Chinese, EUC-KR for Korean, CP866 or Windows-1251 for Russian. Nothing in the file records which one was used.

A flag that marks names as UTF-8 was added to the format later, and plenty of tools and old archives do not set it. When such a ZIP is opened on a computer with a different language setting, the unzip program has to guess, usually guesses its own local code page, and decodes the bytes into the wrong characters. A folder named 日本語 shows up as ô·û{îΩ. The data inside the files is intact; only the labels are misread.

How detection and the manual dropdown work

When the archive is opened, the tool looks at the name bytes of all entries and works out which legacy encoding fits them best. The covered set includes Shift_JIS, GBK and GB18030, Big5, EUC-KR, CP866, Windows-1251, KOI8-R, CP437 and CP850, among others.

Detection is a statistical guess, and short or few names leave little to go on. Two encodings can both produce valid-looking text from the same bytes. That is what the Filename encoding dropdown is for: pick a candidate and the file list re-renders immediately, so you can step through the options until the names read as real words. If you know where the archive came from, start with that region’s encodings.

Saving a corrected ZIP

Fixing the display is half of it. To get an archive that opens correctly everywhere, download everything as a .zip: the tool writes a new ZIP in which every name is stored as proper UTF-8. Current unzip tools on Windows, macOS and Linux read that without guessing. You can also download single files, or use Save to folder in Chrome and Edge to write the tree to disk under the corrected names.

Only the names are changed. File contents are copied as they are, so a text file that was itself written in Shift_JIS or GBK is still in that encoding afterwards.

ZIP only

The encoding fix applies to ZIP archives. Legacy-encoded names in other formats, such as LZH, CAB, old RAR archives or TAR, are not fixed and will still appear garbled. The archive is processed locally and is not uploaded, so private files with private names stay on your computer.

Frequently asked questions

Why are the file names in my ZIP garbled?

The ZIP was created by a tool that stored names in a regional code page without marking which one. Your unzip program decoded those bytes with a different code page, which produces wrong characters, often called mojibake.

How do I fix Japanese file names in a ZIP?

Open the ZIP here. Shift_JIS is normally detected automatically; if not, select it in the Filename encoding dropdown. Then download all files as a .zip to get an archive with UTF-8 names.

The auto-detected encoding is wrong. What now?

Choose a different encoding in the dropdown. The file list updates as you switch, so you can try the likely candidates until the names are readable.

Does this work for RAR, 7z, LZH or TAR archives?

No. The file name encoding fix is for ZIP archives only. Legacy-encoded names in other formats are not corrected.

Are the files themselves changed?

No. Only the stored names are rewritten as UTF-8 in the new ZIP. The content of every file is copied unchanged.