ZeroHour

Search: “qwen3”

9 stories in the last 3d

Show HN: Pelican-bicycle alternatives (updated for 2026)

Hobbyist benchmark re-runs the pelican-bicycle SVG test on six 2026 frontier models, comparing generation time and API cost per image.

A Show HN post re-runs the classic pelican-bicycle and similar SVG generation tests across six 2026 models: GPT-6 Astra, Claude Fable 5.1, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen3.8 Max, and Fugu Ultra v2, recording wall-clock time and cost. It also lists 2025 baseline runs with ten models including Claude Sonnet 4.5, GPT-5.2 Pro, and Qwen3-VL-235B-A22B-Thinking. DeepSeek V4 Pro is consistently cheapest ($0.04-$0.10) while Qwen3.8 Max is slowest, taking up to roughly 17 minutes per generation.

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Nari Labs claims top Coval voice AI benchmark rankings with low-latency, low-cost Qwen3-ASR and Qwen3-TTS inference endpoints.

Nari Labs says its Qwen3-ASR Fast endpoint ranks #1 in Coval's time-to-final-segment latency (p50 44 ms) with 3.6% WER at $0.12/hour, behind only AssemblyAI Universal 3.5 Pro on accuracy. Its Qwen3-TTS Fast ranks #2 in time-to-first-audio (p50 63 ms) and #1 in WER at 3.8%, priced at $10 per 1M characters. The company reports beating the official Qwen3 TTS Flash Realtime endpoint (8.8% WER, 692 ms median TTFA) and Baseten's dedicated endpoint (6.0% WER, 101 ms). Public beta APIs are moving to paid general availability with $20 in credits for existing accounts.

[AINews] Reality Checks on AI News (Yegge shuts down Gas Town, Databricks’ +60% Astra cost)

Latent Space AI news roundup: Steve Yegge shuts down Gas Town, Databricks reports 60% higher coding spend on GPT-6 Astra, OpenAI launches misalignment disclosure framework.

Latent Space's AI News digest for September 15-16, 2026 leads with Steve Yegge shutting down his Gas Town orchestrator despite spending thousands monthly on coding-agent subscriptions. Databricks rolled out GPT-6 Astra to roughly 3,500 engineers, reporting superior long-horizon performance over Opus 5 and Sol 5.6 but a ~60% increase in coding spend. OpenAI published a formal framework for disclosing model misalignment incidents with six case reports, while Microsoft and Google Research released safety papers on 'capability laundering' and the Fuse motive-inference benchmark. Xiaomi shared live RL training telemetry for MiMo-V2.6, estimated at $493k/day for the 1T-class Pro run.

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x higher throughput than GB300 NVL72 and 99% scaling efficiency at 288 GPUs.

In its first MLPerf Inference preview submission, NVIDIA's Vera Rubin NVL72 achieved up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1. A 288-GPU GB300 NVL72 submission across four racks reached 99% scaling efficiency on the DeepSeek-R1 offline benchmark. Software optimizations delivered up to 1.6x gains over v6.0, leveraging TensorRT-LLM, vLLM, Dynamo, disaggregated serving, and NVFP4 precision.

NVIDIA Blog · 17h agoAI industry 2 sources

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

Researchers identify Value Flattening in PPO critics for LLM RL and propose SP^3O sparse value supervision, improving Qwen3-Base training.

The paper uncovers Value Flattening, a failure mode where PPO critic predictions stay flat while true state values estimated from Monte Carlo continuations change sharply, worsening as state spaces grow. The authors attribute it to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states. They propose SP^3O, which supervises value loss on only a few well-separated states per response, consistently improving policies trained on Qwen3-Base across model sizes and evaluation suites.

Hugging Face daily papers · 1d agoAI research

A Zeroth-Order Paradigm for LLM Preference Alignment

ComPO is a zeroth-order preference alignment method using comparison oracles to mitigate likelihood displacement across Mistral, Llama, Gemma, and Qwen3 models.

The paper proposes Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method that extracts directional information from preference pairs with small likelihood margins without directly optimizing a differentiable preference loss. The authors prove convergence guarantees for the offline scheme and performance guarantees for a constrained online variant with reverse-KL control. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 show improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics consistent with mitigating likelihood displacement.

Hugging Face daily papersupdated · 14h agofirst · 1d agoAI research 2 sources

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece improves action tokenization for vision-language-action models via physical rank consistency, reaching 94.8% on LIBERO with Qwen3-VL-4B.

The paper introduces physical rank consistency (PRC), a metric measuring whether tokenization preserves local physical distance rankings of actions after reconstruction. ActionPiece preserves physical action relationships through joint supervision of representation learning and quantization, alongside reconstruction losses. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO, 68.8% on unseen LIBERO-Plus, 71.9% on SimplerEnv, and 51.5% across VLA-Arena L0-L2.

Hugging Face daily papers · 1d agoAI research

Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

Google Research introduced R4T, an RL-trained fan-out pipeline distilled into a 53.9M-parameter diffusion retriever achieving 12x-20x faster query fan-out.

Google Research introduced Retrieve-for-Train (R4T), which trains a fan-out language model with GRPO plus soft PPO regularization, then distills query fan-out into a 53.9M-parameter diffusion transformer that generates all retrieval embeddings in a single non-autoregressive pass. A three-term reward (groundedness 0.6, diversity 0.2 via Vendi Score, alignment 0.2) prevents paraphrastic collapse and reward hacking during training. On the Polyvore dataset, Gemma3-4B R4T-FOLM averaged 49.1 versus 40.9 for Best-of-N, and the diffusion retriever cut fan-out latency from 1.46s to 0.07s at batch size 8, a consistent 12x-20x speedup over autoregressive methods.

MarkTechPost · 2h agoAI research

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

Gavel reads skill-routing signals from a frozen LLM's forward passes with two linear maps, beating retrieval pipelines by up to 21.9 points.

Gavel (Glance And Verdict) shows a frozen agent LLM already contains skill-routing signals in its forward passes, read out via two trained linear maps without loading skill text into context. A glance step scores the full library using mid-layer states and per-skill banks built in one forward pass; a verdict step fuses the model's own likelihood and yes/no judgment as a product of experts. Trained once, it transfers zero-shot to three public benchmarks and SkillTraj (372 simulated agent trajectories); on Qwen3-32B it beats progressive disclosure and retrieve-and-rerank pipelines adding 1.2B-16B external parameters by up to 13.4 points (21.9 mid-rollout).