Alibaba Ovis-OCR2 im Test: Kompaktes 0.8B-Modell für Dokumenten-Parsing

VideoFahd MirzaDemo

Fahd Mirza demonstriert Alibaba Ovis-OCR2, ein Open-Source-Modell mit 0,8 Milliarden Parametern für die strukturierte Markdown-Extraktion aus Dokumenten. Trotz geringer Hardwareanforderungen schlägt das End-to-End-Modell komplexe Pipeline-Systeme in Benchmarks.
Beim Abspielen wird YouTube (youtube-nocookie.com) geladen.

Das Wichtigste

  1. Ovis-OCR2 basiert auf Qwen 3.5 0.8B und wurde mittels SFT, Reinforcement Learning und einer OPD-Technik trainiert.
  2. Das Training stützt sich auf eine zweigleisige Daten-Engine aus realen Dokumenten und synthetisch über LLMs und Playwright erzeugten HTML-Vorlagen.
  3. In Benchmarks wie OmniDocBench und PureDocBench übertrifft das Modell laut Bericht pipelinebasierte OCR-Lösungen bei Tabellen, Formeln und Lesereihenfolge.
  4. Im lokalen Gradio-Test belegt das Modell unter 2 GB VRAM und eignet sich dadurch auch für den CPU-Betrieb.
  5. Praxistests zeigen präzise Erkennung handschriftlicher Texte, komplexer Formeln, dichter historischer Zeitungen und mehrseitiger PDF-Tabellen.
  6. Einschränkung bei Sprachen: Ein Test mit russischem Text scheiterte in Wiederholungsschleifen; der Fokus liegt primär auf Englisch und Chinesisch.

Warum das relevant ist

Klassische Dokumentenverarbeitung verlangt meist mehrstufige Pipelines aus Layout-Analyse, Texterkennung und Nachbereitung. Ovis-OCR2 belegt, dass ultrakompakte multimodale Modelle diese Schritte bei minimalem Hardwarebedarf in einem einzigen Durchlauf bewältigen können.

Einordnung

Das Modell besticht durch eine extrem niedrige Hürde bei der Bereitstellung: Mit unter 2 GB Speicherbedarf lässt es sich problemlos auf handelsüblichen Edge-Geräten oder reinen CPU-Instanzen hosten. Die Praxisbeispiele zeigen Stärken bei komplexen Tabellenstrukturen und Formelsatz. Für rein englisch- und chinesischsprachige Workflows bietet Ovis-OCR2 eine schlanke Alternative zu ressourcenintensiven Modellen.

Transkript

Vollständiges Transkript anzeigen (1.655 Wörter)
Alibaba just open sourced Ovis-OCR2, a tiny but ridiculously powerful document parsing model that punches so well above its weight that it's almost unfair. At just 0.8 billion parameters, this little model takes any document page you throw at it, whether it's packed with complex tables, nasty LaTeX formulas, dense paragraphs, or embedded images and spits out clean, structured Markdown in natural reading order. This model has been built on top of Qwen 3.5 0.8 billion and trained through carefully engineered multistage recipe, combining supervised fine-tuning, reinforcement learning, and a technique called as OPD, which I will describe shortly. This Ovis-OCR2 isn't just another OCR model, if you go through its whole technical report, it's a full end-to-end document intelligence engine that fits anywhere. This is Fahd Mirza and I welcome you to the channel. You can see that we have been covering these Ovis models for quite some time. Let's get right into it, we will install it and then we are going to test it out. I'm going to use this Ubuntu system. I have one GPU card, an NVIDIA RTX A6000 with 48 GB of VRAM. If you're looking to rent a GPU on very good price, you can find the link to Massed Compute with a discount coupon code of 50% for range of GPUs. Also, I have recently started this free weekly AI newsletter, which you can subscribe at fahdmirza.substack.com. What I have done here is, I have gone into their model card, grabbed this code and put a Gradio interface on top of it. So, this is what I'm going to run and I will drop the link to it in video's description. The model card you can also check it out. Let me run this script, which is going to download the model first. The model is being downloaded. It's a small model, as you can see, not really that huge. So, the Gradio demo is running, let me open it in the browser. And this is what it looks like. Let me quickly show you the VRAM consumption. There you go, so just under 2 GB of VRAM, you can easily run it on your own CPU. Okay, so let me upload an image just to convert it into structured Markdown and render it. I'm just going to select something from my local system. So first, let's do this handwritten letter and I am just going to run an OCR on top of it. It is working through it. There is only one page, there you go. So if we check it out, I think so far so good. Even the punctuation marks are there. I'm quickly go through it as I speak. Yep, looks pretty good to me. What do you think? Even the brackets are there. And this is the Markdown. Yeah, looks pretty good, again. Okay, so handwritten looks good. Let's do another one. Let's click on replace. And I'm just going to go with some physics equation, again, handwritten. Let's see if it is able to do that. The speed is quite good as you can already see. Fairly good. And you see the demo is already, you know, fixing it. I'm just very closely checking it out. Let me let it finish and then I will check and let you know. OCR is done and as far as I can see, it has done fairly good. You see even this M bound system, it was able to do it. You see even this, you know, postscript and also all of these indentations are there. Really good. There were even some, you know, different signs in maths and it was able to do that. Let's check the Markdown even better. There you go. Really good stuff. Okay, let's do another one. So I'm just going to do a quick multilingual. Maybe I'll just pick up this Russian and see if it is able to do that Russian one or not. So it is doing something. If you are from Russia, please let us know. This looks, I think it is just repeating it. Doesn't seem like it is multilingual. You see it is just simply repeating it. Okay, I will just stop this and do another one. And I think it is just bilingual English and Chinese. But anyway, let's try out this form. And should be fun to see. There you go, so text frame. And then even this question mark and this axis here. And then it is checking out all the fields and even the text from the fields. Yep, looks pretty good and these buttons are also there. Add help test, that is there. And this is our Markdown. Fairly good. Let's check out this very old newspaper, from New Zealand. I think this is from quite old, I guess. So, there you go. So it has started doing that, you see the even this one, I'll just make it bit bigger, you see. Makes money, money and it has really extracted that, it was quite tiny font. The model quality is really good. I'm just going to scroll here. And I think it is doing fairly well here, really good. I think this is a very good quality model, I just remember the size is very small. And while it generates, let me show you a very interesting bit. Look at this diagram. This tells you the architectural story, which is quite clever. They have built a dual pipeline data engine. One side processes real-world document images through specialized OCR tools, converts them to Markdown via rule-based parsing, and manually quality checks the result. The other side runs the synthetic pipeline, where hard failure cases are mined, fed into multimodal LLM to generate HTML templates diversified by an agent, rendered by playwright, and then paired back as image text training examples with iterative quality control at every step. Both streams feed into a single high-quality training corpus. It's the kind of obsessive data engineering that separates models that benchmark well from models that actually work, and as far as benchmarks are concerned, they are quite wild, I would say. This model doesn't just compete with the giants, it even tops the OmniDocBench leaderboard outright, becoming the first end-to-end model to ever beat the pipeline-based models that previously dominated it across every single dimensions, like text accuracy, formula recognition, table parsing, and reading order. It either leads the pack or sits right at the very top, beating models many times its size, including heavy hitters from Google, Alibaba itself, and various specialized OCR pipelines. For example, on PureDocBench, it claims the highest average score there too. I think this is quite rare, where small doesn't mean compromised, it means efficient, deployable, and quite state-of-the-art, as we are just looking at this result. So huge. Even all the visual regions are there, and then it has extracted it very, very well. Very impressive. Okay, let's do another one. So I'm just checking out this invoice, which has some tabular data, some numbers. Um, so let's see how it goes. So far so good as you can see. The things which we have done already, I think this seems like a trivial task for this. But the table recognition, rows, columns, I think that intelligence is there. There you go. So it is generating fairly good. What happened there? It is still parsing it, there you go. It takes bit of a time, I think the pipeline which we were discussing, it was going through it and going back and forth. And it has completed that table, all the rows are there, and the whole dataset is also there. This one, it was, I think it would look better in the Markdown. There you go. The whole Markdown is there. Fairly good. Let's do maybe a spreadsheet maybe, which are this one is better. Let's do this table one. I think this is even simpler than the previous one. There you go. I just want to check out how exactly it goes with this column. There you go. Yeah, doing pretty well. So you can see that so far, the results are flawless, really good stuff. So, it can even just put the placeholders here and really good. Okay, let me show you maybe a PDF file. And just to change the flavor, I am just going to go with this PDF, and going with the Chinese. There you go. Chinese is even faster. So maybe it's a native language. And I'll just make it bit bigger so that you can also see. By just visually looking at it, I think it has done really good. All the numbers are there, the Chinese characters are there. Of course, I mean, I can't be certain there, but if you are Chinese native speaker or just let us know in the comments, please, what do you think? So far, I think it has done really, really good. And it works with multipage. This is a multipage one and you can see that already. It is doing it page by page. And this is the Markdown if you are wondering. And it has moved on to the page 2. Finally, let's do a quick one on the chart. Let's try this one. And should be fun to see if it is able to get these numbers with percentages. Yep, the numbers are there and even percentages are there. And you can see that document is 5.9, but it has gone with 5.1. 4.9, which was OCR, so OCR and this, this order is not there, but I think it is just going line by line like this, and this is the Markdown. So look, very impressive. Let me know what do you think about this model. If you want to help out the channel, please become a member. Thank you for all the support.

Links und Tools aus diesem Beitrag

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Link:ATH-MaaS

    OvisOCR2: Kompaktes 0,8B-End-to-End-Modell für Dokumenten-Parsing

    ATH-MaaS hat OvisOCR2 veröffentlicht, ein kompaktes multimodales Modell mit 0,8 Milliarden Parametern zum Dokumenten-Parsing auf Seitenebene. Das Modell basiert auf Qwen3.5-0.8B und wandelt Dokumentbilder direkt in strukturiertes Markdown mit Tabellen, LaTeX-Formeln und Bild-Bounding-Boxes um.

    KI & AI· Tool

  • Link:docling.ai

    Docling: Open-Source-Dokumentenverarbeitung für KI und RAG

    Docling ist ein quelloffenes Werkzeug der LF AI & Data Foundation, das komplexe Dokumente wie PDFs, Office-Dateien und Bilder in strukturierte Daten für Sprachmodelle umwandelt. Es arbeitet lokal und offline, erkennt Layouts sowie Tabellenstrukturen und stellt Schnittstellen für Python, CLI, REST-APIs und MCP bereit.

    KI & AI· Tool

  • Repository:docling-project/docling

    Docling: Dokumenten-Parsing für Generative AI

    Docling ist ein Open-Source-Python-Werkzeug, das Dokumente aus zahlreichen Formaten wie PDF, Office-Dateien, Audio und HTML für den Einsatz in Generative-AI-Pipelines aufbereitet.

    66.025SternePython

    KI & AI· Tool

  • Link:Qwen

    Qwen3.8-27B: Multimodales 27B-Modell mit Hybrid-Architektur und Denkmodus

    Das Qwen-Team hat mit Qwen3.8-27B ein kompaktes, dichtes Vision-Language-Modell unter der Apache-2.0-Lizenz veröffentlicht. Das 27-Milliarden-Parameter-Modell kombiniert Gated DeltaNet mit regulärer Attention, versteht Bilder sowie Videos und bietet flexible Steuerungsmöglichkeiten für Denkprozesse.

    KI & AI· Sammlung

  • X-Post:Nicolas Camara

    Firecrawl veröffentlicht pdf-inspector: Schnelle PDF-Klassifizierung und Markdown-Extraktion in Rust

    Firecrawl hat pdf-inspector vorgestellt, eine quelloffene Rust-Bibliothek zur raschen Erkennung von PDF-Typen und direkten Markdown-Konvertierung ohne zwingendes OCR. Das Tool zielt darauf ab, KI-Agenten und Parsing-Pipelines zu beschleunigen, indem rein textbasierte Dokumente innerhalb von Millisekunden lokal extrahiert werden.

    20.611Lesezeichen2,7 Mio.Aufrufe

    KI & AI· Tool

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.