The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · Amazon · ALoDLM · Qwen3 · Hugging Face

Amazon releases diffusion language model it says beats Qwen3 at 2.7x speed

Amazon's ALoDLM spends extra compute only on hard tokens. The authors say the 8B model scores 80.3 across eleven benchmarks against 78.5 for Qwen3-8B, and runs about 2.7 times faster on GSM8K.

Amazon researchers have released ALoDLM, a diffusion language model that tries to close the quality gap between parallel-decoding models and ordinary autoregressive ones. The paper took 33 upvotes on Hugging Face's daily papers page. Weights for 1.7-billion and 8-billion-parameter versions are on Hugging Face, with code on GitHub. Nobody outside Amazon has reported testing them yet.

The authors argue that diffusion models waste effort by giving every unknown token the same depth of computation, although some are easy to predict and others are not. ALoDLM instead loops extra refinement passes on unresolved tokens, while tokens that are ready get committed as fixed context. The schedule is learned during training rather than hand-labelled.

On eleven benchmarks the authors report an average of 80.3 for the 8B model against 78.5 for Qwen3-8B, and 65.5 against 63.8 at 1.7B. The authors say it outperforms every diffusion model they evaluated. On GSM8K they report about 2.7 times the throughput of Qwen3-8B served with vLLM, and 612.4 tokens per second on one Nvidia B200. These are the authors' measurements.

The repository carries a non-commercial license, CC BY-NC 4.0, for the original work, so companies cannot build products on it as released. It had two stars and one fork when this paper looked, so it reads as a research drop rather than a community favourite. Independent reproductions of the benchmark numbers have not appeared.

Sources 3 sources

  1. Source Amazon researchers
  2. Source ALoDLM project page
  3. Source ALoDLM repository