August 2026
SlopFinder: Crowdsourcing AI Slop Classification
Building an open dataset to study what humans actually consider AI slop. SlopFinder presents random AI-generated text samples with a single rating control, streaming crowdsourced human evaluations directly to Hugging Face.
Read article →
August 2026
PixelModel v6: MMDiT, REPA, and the Capacity Bet
Six real architecture changes and nearly 4x the capacity: MMDiT joint attention, T5 conditioning, 2D RoPE, QK-norm + SwiGLU, and REPA. FID drops to 23.62 at cfg 3.0, beating v5's 34.06 by over ten points at every guidance scale tested.
Read article →
July 2026
PixelModel v5: The Bottleneck Was Data, Not Architecture
Same DiT as v4, 36x more training data: FID drops 39.54 → 34.06 and CLIP score climbs to 29.27, within 0.78 of the real-photo ceiling. Nothing in the network changed.
Read article →
July 2026
VoxelModel v1: text to 3D
Extending implicit coordinate conditioning to (x, y, z) space to generate resolution-independent 3D shapes with 148k parameters.
Read article →
July 2026
PixelModel v4: Cross-Attention, Text Transformers, and Native 256px
v4 replaces bag-of-words mean-pooling with a 2-layer text Transformer and spatial cross-attention, breaking FID 320 at native 256px.
Read article →
July 2026
PixelModel v2 and v3: Scaling Up, Then Doing It Right
v2: scaled up, v3: a different architecture
Read article →
July 2026
The 7-2026 generation: dual-mode benchmarks
All three tiers remade: 300 balanced items each, one unified schema scored by generation and log-likelihood in a single run, machine-verified Math, blind cross-checked answers, full per-item provenance. 6-2026 stays frozen; the leaderboard now defaults to the latest generation.
Read article →
July 2026
Dataset cleanup v2
Deduped Easy (269→238) and Mid (155→143), fixed a broken barber-paradox item, unified the Effortless taxonomy, and corrected two metric labels. Existing leaderboard scores are flagged stale pending re-eval.
Read article →
July 2026
One Script, Every Benchmark
The universal evaluator is live: download script.py from the leaderboard Space, point it at any HF model, get every benchmark scored with full category/subcategory breakdowns — and a ready-to-paste models.json entry.
Read article →
July 2026
PixelModel v1: The Model Is a Thumbnail Now
8.5× fewer parameters, better FID and CLIP Score, any-resolution output — and all 23,747 weights still live in one PNG, now the size of a thumbnail.
Read article →
July 2026
Latent Reasoning Directions (LRD)
the model is not being taught how to reason step by step. Instead, it learns which reasoning mode to activate before solving the problem. a lightweight direction module to predict a latent reasoning vector using a regression-style latent-space objective.
Read article →
July 2026
new drop, less slop
We hear you. Every time. We decided its time to talk about it.
Read article →
July 2026
bench-AGI
Finally, a great test for AI.
Read article →
June 2026
bench-mid-6-2026 · Qwen/Qwen2.5-0.5B benchmarked
Evaluation of Qwen/Qwen2.5-0.5B on bench-labs/bench-mid-6-2026 using lm-eval multiple-choice
loglikelihood scoring (CPU, bf16).
Read article →
June 2026
bench-easy-6-2026, we ran the benchmark.
It was good, but not perfect. An easy-tier benchmark. Bench-labs/bench-easy-6-2026 tested on Qwen 1.5B
Read article →
June 2026
Lm Eval Glint 1.3
The code is present.
Read article →
June 2026
We are looking for people
We seek colaborators to utilize our benchmarks and work with us. More information at spaces/Bench-labs/Members
Read article →
New people have joined us!!
Open the article to see who's new..
Read article →
June 2026
bench-effortless-6-2026
An effortless-tier benchmark, 240 rows total, see at Bench-labs/bench-effortless-6-2026
Read article →
June 2026
Some Q&A - benchmarks
A small post
Read article →
June 2026
PixelModel: When the Weights Are the Image
A neural network where the weights are literally encoded in the pixels of a PNG image.
Read article →
June 2026
Introducing Bench Labs
Why we started Bench Labs, what we mean by
generalization, and what we're planning to build.
Read article →
June 2026
Formatting Debug
Throwaway post exercising every styled element.
Read article →