Supported file formats
Overview
Smartcat supports a wide variety of file formats for translation and localization. When you upload a file, Smartcat detects the format automatically and routes it to the appropriate processing engine.
Drag and drop your files or select them when creating a new project or translation task. If a file format is not supported, an error message appears and the file is not uploaded.
Supported file formats
Microsoft Office
-
DOC / DOCX
-
XLS / XLSX / XLSM
-
PPT / PPTX / PPS / PPSX / POT / POTX
📌 Microsoft Office formats are proprietary and complex. Some post-processing may be required on translated documents.
Open Office
-
ODT
-
ODP
Text and rich text
-
TXT
-
RTF
Data and Spreadsheets
- CSV
Markdown
- MD / MKD / MDWN / MDOWN / MDTXT / MDTEXT / MARKDOWN
Hypertext
-
HTML / HTM
-
PHP
Bilingual interchange formats
-
XLIFF (XLF) 1.2 and 2.0
-
SDLXLIFF
-
MQXLIFF
-
XLIFF files from Articulate Rise 360
-
XLIFF files from Articulate Storyline
-
XLIFF files from Easygenerator
-
PO / POT
-
TTX
Desktop publishing
-
PDF
-
IDML (Adobe InDesign)
-
INX
-
MIF (FrameMaker)
📌 PDF files are processed using OCR. See the Images (OCR) section for supported input formats.
Technical writing
-
DITA XML
-
DITAMAP
-
HELP+MANUAL XML
Localization
-
XML
-
TTML
-
Android XML
-
RESX
-
LOCJSON
-
JSON
-
TJSON
-
YML / YAML
-
INC
-
INX
-
STRINGS (Apple)
-
STRINGSDICT (Apple)
-
XCSTRINGS (Apple Xcode string catalog)
-
PROPERTIES (Java)
Learning content
-
Articulate Rise SCORM course (ZIP)
-
Articulate Storyline course (ZIP)
-
Articulate Storyline native file (STORY)
Video
-
MP4
-
MPEG / MPG / M2V
-
AVI
-
MOV / QT
-
MKV
-
M4V
-
FLV
-
3GP / 3G2
-
OGV
-
TS
-
WMV
-
VOB
Audio
-
MP3
-
MP2 / M2A
-
M4A
-
AAC
-
OGG
-
FLAC
-
WMA
Subtitle
-
SRT
-
VTT
Images (OCR)
The following image and document formats are supported via OCR processing:
-
JPG / JPEG / JFIF
-
TIF / TIFF
-
PNG
-
BMP
-
GIF
-
DCX / PCX
-
JP2 / JPC
-
DJVU / DJV
-
JB2
-
PDF (image-only)
-
AI (Adobe Illustrator)
Packages
-
SDLPPX / SDLRPX
-
WSXZ
-
ZIP
📌 ZIP packages are used for IDML, DITA, Articulate Rise SCORM, Articulate Storyline courses, and other packaged content. Smartcat detects the package type automatically.
Unsupported formats
- FM (FrameMaker binary)
Blocked file extensions
To protect against security risks, Smartcat blocks upload of files with the following extensions. These cannot be uploaded regardless of content:
ade, adp, apk, app, appx, appxbundle, asp, aspx, asx, bas, bat, cab, cer, chm, cmd, cnt, com, cpl, crt, csh, der, diagcab, dll, dmg, exe, fxp, gadget, grp, hlp, hpj, hta, htc, inf, ins, iso, isp, its, jar, jnlp, js, jse, ksh, lib, lnk, mad, maf, mag, mam, maq, mar, mas, mat, mau, mav, maw, mcf, mda, mdb, mde, mdt, mdw, mdz, msc, msh, msh1, msh1xml, msh2, msh2xml, mshxml, msi, msix, msixbundle, msp, mst, msu, nsh, ops, osd, pcd, pif, pl, plg, prf, prg, printerexport, ps1, ps1xml, ps2, ps2xml, psc1, psc2, psd1, psdm1, pst, py, pyc, pyo, pyw, pyz, pyzw, reg, scf, scr, sct, sh, shb, shs, sys, theme, tmp, url, vb, vbe, vbp, vbs, vhd, vhdx, vsmacros, vsw, vxd, webpnp, website, ws, wsc, wsf, wsh, xbap, xll, xnk
⚠️ Note: js appears in the blocked extensions list above. If you need to translate JavaScript localization files, export them as JSON or another supported localization format.
Format-specific limitations and best practices
This section addresses common questions about specific file format handling.
Microsoft Office files
| Format | Supported | Known Limitations |
|---|---|---|
| DOCX | ✅ | Complex formatting (nested tables, text boxes) may require post-processing |
| Bilingual DOCX (parallel columns) | ⚠️ Partial | Export only — cannot import edited bilingual DOCX back into Smartcat |
| XLSX | ✅ | Use MultilingualExcel for multi-language files with source/target columns |
| XLSM (macro-enabled) | ✅ | Macros are preserved but not translated |
PDF files
| Scenario | Supported | Notes |
|---|---|---|
| Native PDF (text-based) | ✅ | Full support with text extraction |
| Scanned PDF | ✅ | OCR automatically extracts text from scanned pages |
| Image-heavy PDF | ✅ | OCR processes embedded images containing text |
| PDF with complex layouts | ✅ | Layout preservation depends on document complexity |
📌 For best results with scanned PDFs, ensure the source document has clear, high-resolution text.
E-Learning formats
| Format | Supported | Notes |
|---|---|---|
| SCORM 1.2 | ✅ | Full support via Learning Content Agent |
| SCORM 2004 (all editions) | ✅ | Full support via Learning Content Agent |
| Articulate Rise | ✅ | Native support — upload SCORM ZIP package |
| Articulate Storyline | 🟡 On roadmap | Currently on short-term roadmap |
| Articulate Storyline (XLIFF export) | ✅ | Supported via XLIFF file upload |
| Articulate Storyline (.STORY native) | ✅ | Native file support available |
| xAPI (Tin Can) | ❌ | Not currently supported |
Localization formats
| Format | Supported | Notes |
|---|---|---|
| JSON | ✅ | Standard JSON key-value pairs supported |
| JSON (nested objects) | ✅ | Nested structures supported with proper formatting |
| XLIFF 1.2 | ✅ | Full support including translation state attributes |
| XLIFF 2.0 | ✅ | Full support |
| YAML | ✅ | Standard YAML localization files supported |
Desktop publishing
| Format | Supported | Notes |
|---|---|---|
| IDML (InDesign) | ✅ | Original formatting preserved; tag settings available for optimization |
📌 For IDML files, review the tag settings before translation to optimize segment handling. See: https://help.smartcat.com/idml-file-tag-settings/
FAQ — Format-specific questions
Q: Can Smartcat handle bilingual DOCX files with parallel columns?
A: Smartcat can export bilingual DOCX files with parallel source/target columns. However, importing edited bilingual DOCX files back into Smartcat is not currently supported.
Q: Do we support JSON Objects (nested structures)?
A: Yes, JSON files with nested object structures are supported. Ensure your JSON is properly formatted with standard key-value pairs.
Q: What happens with scanned PDFs in the PDF Agent?
A: Scanned PDFs are automatically processed using OCR (Optical Character Recognition) to extract text for translation. For best results, use high-resolution scans with clear text.
Q: Can we translate IDML files without quality degradation?
A: Yes, IDML files preserve original formatting. Use the IDML tag settings to optimize segment handling and simplify text formatting.
Q: Do we support xAPI (not just SCORM)?
A: xAPI (Tin Can) is not currently supported. SCORM 1.2 and SCORM 2004 (all editions) are fully supported.
Q: Can we handle Articulate Storyline audio files?
A: Articulate Storyline native support is on the short-term roadmap. Currently, you can export Storyline content as XLIFF for translation.
Still need help?
Our support team responds within one business day.