ZeroHour

Search: “ocr”

10 stories

How Far Can Synthetic Data Take Thai OCR?

Synthetic-only training adapts PaddleOCR-VL into Wayu-Paxa-OCR-Zero, cutting Thai printed-page CER from 6.64% to 1.24% without real labels.

The study disentangles which factors of synthetic OCR data transfer to real Thai documents, finding typeface diversity, 2D structure, and real handwriting glyphs matter most. Using 45,723 synthetic pages, the authors adapt the 0.9B-parameter PaddleOCR-VL-1.6 into Wayu-Paxa-OCR-Zero, reducing median CER from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting. The model outperforms Typhoon OCR v1 7B on all five evaluation sets.

Hugging Face daily papers · 14d agoAI research

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

Reducto launched r-1, a single-pass document parsing model claiming 20% error reduction over its legacy agentic pipeline, priced at 1 cent per page.

Reducto announced r-1, the first model in a new parsing family that replaces multi-stage agentic OCR with one full-page pass handling text, tables, figures, layout, formatting, and grounding with page-relative bounding boxes. The company reports a 20% error reduction measured against its own legacy agentic pipelines, plus vendor-run wins over Amazon Textract and Azure Document Intelligence on complex documents. Pricing is a flat 1 cent per page versus 3-6 cents for legacy models; r-1 is available in preview via the V3 Parse API with no open weights.

MarkTechPost · 9d agoModel release1

TaichuAI/ZDTaichu5.0-9B — new model trending #30 on Hugging Face

TaichuAI released ZDTaichu5.0-9B, an open multimodal VLM built on Qwen3.5-9B targeting spatial reasoning, embodied AI, and agentic tool use.

TaichuAI released ZDTaichu5.0-9B, a multimodal vision-language model pairing a Qwen3.5-9B language decoder with a C-RADIOv4-H vision encoder, supporting text, single/multiple images, and video with any-resolution input and a 128K-token context. It introduces Entropy-Gated Adaptive Recurrent Reasoning, which allocates extra latent refinement steps to harder tokens. Reported benchmarks include 87.7 on TAU2-Bench, 71.4 on Claw-Eval, 93.7 on IFEval, 48 on ERQA, and 56 on RoboSpatial, leading compared 10B-scale open VLMs on agent and spatial tasks. The weights are available on Hugging Face, GitHub, and ModelScope, where it is trending at #30.

Hugging Face trending models · 13d agoModel release1

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1 Flash is a 552B-parameter multimodal MoE model with 1M-token context achieving 4x KV cache compression for long-horizon agent workloads.

A detailed analysis of the DeepSeek-V4.1 Flash technical report describes a 552B-parameter multimodal mixture-of-experts model supporting contexts up to 1 million tokens. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters during prefill and 16B during decode, and reportedly delivers about 420 tokens/s. Joint optimization of architecture (CSA2 cross-layer compression), FP4 KV cache precision, and deployment strategy cuts runtime KV cache to roughly 1/4 and persistent KV cache to about 1/8 of DeepSeek-V4-Flash at the same sequence length, targeting storage and bandwidth bottlenecks in long-horizon agent serving. The author notes all DeepSeek-V4 Pro models were taken offline following the release.

Tables Decoded: DELTA for Structure, TARQA for Understanding

DELTA extracts tables into compact OTSL text and TARQA fine-tunes LLMs on it, beating VLM baselines on table QA.

DELTA separates physical structure recognition, logical structure recognition, and OCR to output tables in Optimised Table Structure Language (OTSL), a compact unified format encoding cell arrangements and content. It achieves TEDS-Structure scores comparable to state-of-the-art methods across FinTabNet, PubTabNet, and PubTables-1M, with robustness tested on a curated Hindi benchmark, TORQUE. TARQA, an LLM fine-tuned on OTSL sequences, gains 9.3 percentage points on WTQ TabQA and 9.2 points on FinTabNetQA TabVQA; code, models, and the benchmark are released on GitHub.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Tell HN: iOS 27 does not allow Apple Intelligence to be disabled

Hacker News discussion argues Apple Intelligence is a marketing umbrella for disparate long-standing ML features, explaining why iOS 27 lacks a disable switch.

A 'Tell HN' post reports that iOS 27 does not allow users to fully disable Apple Intelligence. Discussion replies argue 'Apple Intelligence' is a marketing term covering disparate features rather than a single product. Commenters cite long-standing local-model features such as image OCR, battery-life prediction, and the Latent Semantic Mapping text classification framework available since 2007.

When the prompt becomes the payload: A practical pen-testing guide for GenAI, LLM and RAG applications

CSO Online publishes a practical penetration-testing guide for GenAI, LLM, and RAG applications, covering prompt injection, retrieval poisoning, and tenant isolation testing.

The guide frames LLM applications as attack graphs spanning prompts, retrieval layers, vector stores, tools, identities, and downstream APIs, arguing that conventional web testing misses instruction-vs-data channel risks. It builds on OWASP prompt injection guidance (direct vs. indirect injection) and NIST's 2025 adversarial machine-learning taxonomy, noting that RAG and fine-tuning do not remove injection risk. Recommended practices include documenting trust transitions across components, using canaries and synthetic records to avoid test side effects, running multi-turn and obfuscated injection campaigns, and verifying chains from poisoned documents to observable state changes. It also details testing RAG pipelines via controlled document poisoning across metadata, OCR layers, and code comments, plus cross-tenant isolation checks on retrieved document IDs.

CSO Online · 8d agoAI safety & security1