Cookies for analytics and advertising
We use cookies for analytics and advertising, both sent to Google. Refusing changes nothing you can see.Read the privacy page
Converting AVIF to TXT reads the words out of a picture and hands them back as plain text you can paste — and on this site it takes two steps, because the recogniser has no AVIF decoder. Convert the AVIF to PNG first, then read the text out of the PNG. Both steps run on your own machine.
Up to 100 files at once. Mixed formats are fine.
They convert one after another and download together as a ZIP.
AVIF to TXT
This is the first thing to know and it belongs at the top rather than in a footnote. The recogniser is Tesseract compiled to WebAssembly, and it is handed the file exactly as you dropped it — no browser decoding happens in front of it, deliberately, because pushing an image through a canvas first would flatten transparency to black and a scan of white paper is precisely where that ruins the result.
What that means is that the only image readers in play are the ones built into the recognition engine, and there are six of them: BMP, JPEG, PNG, PBM, WebP and non-animated GIF. AVIF is not among them. There is no AVIF reader inside the WebAssembly core this site serves, so dropping an AVIF here does not produce bad text — it produces no text at all.
The route that works is two steps and it costs you a file. Convert the AVIF to PNG on this site, then read the text out of the PNG. PNG is one of the six the recogniser opens, it is lossless, and it is the right choice for the intermediate file for exactly that reason: the pixels the recogniser sees are the pixels the AVIF decoder produced, with nothing added on the way through.
JPEG would also be read, and it is the wrong choice here. Saving through JPEG re-compresses letterforms that AVIF has already softened once, and the second pass lands on the same thin strokes the recogniser is trying to match. A PNG in the middle is larger on disk and better in every way that affects the reading — and it is a file you can throw away as soon as the text is out.
Nobody exports AVIF from a scanner or a phone gallery. It is a web delivery format, published in 2019 by the Alliance for Open Media, and it reaches you because a site with a modern image pipeline chose it — most often because your browser advertised that it could read the format and the server picked the smallest file it had.
That means the material arriving on this page has a particular character. It is what people save off pages: a chart, a comparison table, a slide from a deck, a quotation set as an image, a price list, a document somebody scanned and published as a picture instead of as text. In every case the words are visible and unselectable, and recognising them is the only route short of retyping.
It also means there is no shortcut hiding inside the file. Some formats carry a text layer alongside the picture — a PDF from a word processor does, and selecting from it costs nothing. An AVIF holds compressed pixels and a little metadata, so there is nothing to extract and everything to recognise. That distinction decides how much you have to check afterwards.
Different codecs fail differently, and how a codec fails decides how recognition fails. JPEG breaks down into blocks and ringing — bright halos around hard edges — which leaves letterforms noisy but still sharply bounded. AVIF, being a still frame of a video codec, tends to smooth instead: it flattens what it decides is not worth describing.
Smoothing is kinder to look at and harder to read. Thin strokes lose their edges against the paper, counters inside letters close up, and adjacent characters merge into one shape the recogniser then guesses at. That is why an AVIF that looks perfectly clean at page size can read badly at eight-point type, and why judging whether the file is good enough means zooming in on the small text rather than glancing at the whole image.
A modern site rarely holds one copy of an image. It holds several sizes and formats and hands over whichever fits the layout and the browser, which is why two people saving the same picture from the same page can end up with different files. What you have is the variant that suited the width of your window at the moment you right-clicked.
That is worth thirty seconds before converting, because a larger variant is worth more than any setting on this page. Opening the image in its own tab often yields a version two or three times the size, and two or three times the pixels per character is the difference between a mangled reading and a clean one. Where the page links a full-size version behind the thumbnail, use that instead of what the layout gave you. Enlarging the file you already have does the opposite of helping: it adds pixels without adding information and softens the letterforms further, so the recogniser gets a bigger version of the same ambiguity.
This pair has a problem the other recognition pages do not: the reader frequently cannot see the source. AVIF is read by every current browser and by a good deal less outside one, so double-clicking the file may do nothing at all, and checking a recognised word against the original is exactly what you need to do most.
The fix is to drag the AVIF into a new browser tab. It will display, it will zoom, and you can put it beside the recognised text. And the two-step route does not depend on that working: the conversion to PNG uses an AVIF decoder compiled to WebAssembly and shipped with the page, not the one in your browser or your operating system, so a machine that shows you a blank thumbnail is not in the way.
The output is a stream of characters in the order they were read, and nothing else. For a paragraph of prose that loses very little. For a chart it means you get the axis labels, the series names and the numbers with no indication of which belonged to which, because the arrangement that made it a chart was spatial and plain text has no space.
For a table it is the same problem with sharper consequences: the cell values arrive in reading order and the columns are gone, so rebuilding it takes about as long as reading it did. That is worth knowing before starting on a screenshot of a spreadsheet. Where the underlying data exists anywhere — a published file, an export, the page the picture came from — finding it beats recognising a picture of it.
Recognition matches shapes to characters without understanding any of them, and it has no mechanism for flagging uncertainty. The output reads as fluent and complete whether every character was right or a quarter of them were guesses, which is precisely what makes it dangerous.
The material saved from web pages is heavy on exactly the content where a single wrong character matters: prices, percentages, dates, statistics, reference numbers, model codes. A misread digit inside a figure you then quote is the error most likely to survive a proofread and least likely to be forgiven. Read the numbers back against the image, and pay particular attention to the digit pairs that share a silhouette once the strokes have been softened.
The pragmatic habit is to treat the output as a draft that saved you the typing rather than as a transcript. Words in prose are self-correcting, because a wrong one usually reads as wrong. A wrong digit reads as a number.
The recogniser works inside one language’s letterforms and vocabulary. Given the wrong one it does not report a failure — it returns the closest match available in a set that cannot contain the right answer, which is why the classic symptom is a page of plausible words that mean nothing.
English, German, French and Spanish are the options here, and the choice matters more than it looks on material saved from the web, where a page you were reading in translation or a screenshot from another country is routinely not in the language your machine defaults to. If a result is strange in a way you cannot account for, check this setting before checking anything else.
Recognition is not always the cheaper path and it is worth saying so on the page that sells it. A short caption, a heading, a handful of figures from a chart, a single line of a diagram: by the time you have converted, read the output and checked it against the picture, you could have typed it and known it was right.
Where it pays is volume and density — a full page of body text, a long table, a folder of scanned documents, anything where retyping is measured in tens of minutes. The threshold is lower than people expect for clean type and higher than they expect for a small, heavily compressed web image, which is the kind this page most often gets.
Nothing, at either step. The AVIF decoder writes the PNG on your own processor and Tesseract reads it in the same tab, both as WebAssembly. Neither file is uploaded, there is no account, no queue and no daily allowance, and the free ceiling is 100 MB per file — far past anything a website will have served you.
That matters more for this pair than the ordinary privacy argument suggests. A large share of what gets saved as an image and read back is material somebody would rather not hand to a third party: an internal dashboard, a page behind a login, a document circulated in a private group, a screenshot from a system with a name on it. Here the file stays where it already is.
The output is a plain text file — Unicode characters and line breaks, with no fonts, no sizes, no bold and no structure. Notepad, TextEdit and Visual Studio Code all open it, as does every editor, every browser and every scripting language, which is the point of the format and has been since the idea of a character set was standardised in the sixties.
That plainness is what makes the result useful downstream. Text this bare pastes into anything, is searchable immediately, diffs cleanly, and can be fed to a script without a parser. If you need the words back in a laid-out document, take the text into a word processor and build the layout there — reconstructing it is a task for a person who knows what the document meant, not for a recogniser reading shapes.
| AVIF | TXT | |
|---|---|---|
| Full name | AV1 Image File Format | Plain Text |
| File extension | .avif | .txt, .text, .log |
| Media type | image/avif | text/plain |
| Compression | Either, depending on the setting | — |
| First published | 2019 | 1963 |
| Published by | Alliance for Open Media | — |
| Specification | AV1 Image File Format | Unicode |
| Licensing | Open standard | Open standard |
| Standing today | Current | Current |
| Bit depth | 12 | — |
| Colour it can describe | RGB, YCbCr, wide gamut | — |
| Largest image | 65,536 px per side | — |
| Opens in a browser | Current browsers | Every browser |
| Considered instead | WebP, JXL, JPG | MD, RTF |
TXT is a working format and AVIF is a finished one. What comes back is editable text and objects rather than a picture of a page, which is usually the reason for the conversion and also where its limits are.
TXT opens in every current browser. AVIF has narrower browser support than that. If the file is going onto a web page or into a form, that is usually the whole reason for the conversion.
The usual programs do not overlap: AVIF opens in GIMP, Squoosh and ImageMagick, TXT in Notepad, TextEdit and Visual Studio Code — so whoever receives the result needs something from the second list.
The two are aimed at different work: AVIF at the web and handing a finished file over, TXT at moving data between programs and archiving. That is worth weighing before converting, because the reason one exists is usually the reason the other is awkward.
AVIF is Alliance for Open Media's format, published in 2019. It records 12 bits per channel.
TXT dates from 1963, specified as Unicode. Notepad, TextEdit and Visual Studio Code all read it.
TXT was published in 1963 and AVIF in 2019. The older one is generally the safer file to hand to somebody; the newer one usually does the job in fewer bytes.
No. This conversion runs entirely inside your browser, so the file never leaves your device. You can confirm it yourself: open the network tab of your browser's developer tools and convert something. You will see the page load, plus the analytics and advertising the site is paid for with — and nothing carrying your file. The engine behind this particular pair is Tesseract, the open-source text recognition engine; your browser fetches it once and caches it.
Yes, but not in one step. The recogniser reads the file itself and its decoders cover BMP, JPEG, PNG, PBM, WebP and GIF, so an AVIF dropped straight on it produces nothing. Convert the AVIF to PNG here first, then read the text out of the PNG.
Two reasons that both come from the web. The image was served at the size the layout needed rather than at reading resolution, and AVIF compression tends to soften small type rather than sharpen it. Thin strokes blur into the background and letters merge.
Drag the file into a new browser tab. Every current browser displays AVIF even where the operating system does not, so you can zoom in on the picture in one tab and read the recognised text beside it.
No. The output is plain text in reading order, so a table comes back as a stream of cell values with no columns. Useful for capturing the figures, not for rebuilding the table — where the underlying data exists somewhere, finding it is faster.
Set the language before converting. Read against the wrong one, the recogniser returns the closest match inside a vocabulary that cannot contain the right answer, so the output is confident nonsense rather than an error. English, German, French and Spanish are available.
No. Both steps run inside your browser tab on your own processor: a WebAssembly AVIF decoder makes the PNG, and the recogniser reads that. Nothing is uploaded, there is no account and no daily allowance, and a folder of images converts in one pass.
The claims this page makes about AVIF and TXT are checkable, and these are the documents that settle them.