Bench-labs Blog

Peak and see

August 2026

SlopFinder: Crowdsourcing AI Slop Classification

Building an open dataset to study what humans actually consider AI slop. SlopFinder presents random AI-generated text samples with a single rating control, streaming crowdsourced human evaluations directly to Hugging Face.

Read article →
August 2026

PixelModel v6: MMDiT, REPA, and the Capacity Bet

Six real architecture changes and nearly 4x the capacity: MMDiT joint attention, T5 conditioning, 2D RoPE, QK-norm + SwiGLU, and REPA. FID drops to 23.62 at cfg 3.0, beating v5's 34.06 by over ten points at every guidance scale tested.

Read article →
July 2026

PixelModel v5: The Bottleneck Was Data, Not Architecture

Same DiT as v4, 36x more training data: FID drops 39.54 → 34.06 and CLIP score climbs to 29.27, within 0.78 of the real-photo ceiling. Nothing in the network changed.

Read article →
July 2026

VoxelModel v1: text to 3D

Extending implicit coordinate conditioning to (x, y, z) space to generate resolution-independent 3D shapes with 148k parameters.

Read article →
July 2026

PixelModel v4: Cross-Attention, Text Transformers, and Native 256px

v4 replaces bag-of-words mean-pooling with a 2-layer text Transformer and spatial cross-attention, breaking FID 320 at native 256px.

Read article →
July 2026

PixelModel v2 and v3: Scaling Up, Then Doing It Right

v2: scaled up, v3: a different architecture

Read article →
July 2026

The 7-2026 generation: dual-mode benchmarks

All three tiers remade: 300 balanced items each, one unified schema scored by generation and log-likelihood in a single run, machine-verified Math, blind cross-checked answers, full per-item provenance. 6-2026 stays frozen; the leaderboard now defaults to the latest generation.

Read article →
July 2026

Dataset cleanup v2

Deduped Easy (269→238) and Mid (155→143), fixed a broken barber-paradox item, unified the Effortless taxonomy, and corrected two metric labels. Existing leaderboard scores are flagged stale pending re-eval.

Read article →
July 2026

One Script, Every Benchmark

The universal evaluator is live: download script.py from the leaderboard Space, point it at any HF model, get every benchmark scored with full category/subcategory breakdowns — and a ready-to-paste models.json entry.

Read article →
July 2026

PixelModel v1: The Model Is a Thumbnail Now

8.5× fewer parameters, better FID and CLIP Score, any-resolution output — and all 23,747 weights still live in one PNG, now the size of a thumbnail.

Read article →
July 2026

Latent Reasoning Directions (LRD)

the model is not being taught how to reason step by step. Instead, it learns which reasoning mode to activate before solving the problem. a lightweight direction module to predict a latent reasoning vector using a regression-style latent-space objective.

Read article →
July 2026

new drop, less slop

We hear you. Every time. We decided its time to talk about it.

Read article →
June 2026

bench-mid-6-2026 · Qwen/Qwen2.5-0.5B benchmarked

Evaluation of Qwen/Qwen2.5-0.5B on bench-labs/bench-mid-6-2026 using lm-eval multiple-choice loglikelihood scoring (CPU, bf16).

Read article →
June 2026

bench-easy-6-2026, we ran the benchmark.

It was good, but not perfect. An easy-tier benchmark. Bench-labs/bench-easy-6-2026 tested on Qwen 1.5B

Read article →
June 2026

We are looking for people

We seek colaborators to utilize our benchmarks and work with us. More information at spaces/Bench-labs/Members

Read article →

New people have joined us!!

Open the article to see who's new..

Read article →
June 2026

bench-effortless-6-2026

An effortless-tier benchmark, 240 rows total, see at Bench-labs/bench-effortless-6-2026

Read article →
June 2026

PixelModel: When the Weights Are the Image

A neural network where the weights are literally encoded in the pixels of a PNG image.

Read article →
June 2026

Introducing Bench Labs

Why we started Bench Labs, what we mean by generalization, and what we're planning to build.

Read article →
June 2026

Formatting Debug

Throwaway post exercising every styled element.

Read article →

All rights reserved to @bench-labs