← Back
June 2026 · Benchmark

bench-mid-6-2026 — Qwen/Qwen2.5-0.5B

Evaluation of Qwen/Qwen2.5-0.5B on bench-labs/bench-mid-6-2026 using lm-eval multiple-choice loglikelihood scoring (CPU, bf16).

Headline Results

Metric Value StdErr
acc0.58060.0398
acc_norm0.60000.0395
soft_score0.58710.0394
soft_score_norm0.61030.0388

Per-Category Breakdown

Category N acc acc_norm soft_score soft_score_norm
Commonsense-causality5 1.0001.000 1.0001.000
Commonsense-reasoning10 0.7000.600 0.7100.620
Commonsense-simulation10 0.3000.500 0.3000.500
Knowledge-basic8 1.0001.000 1.0001.000
Knowledge-definitions10 0.7000.900 0.7000.900
Language-comprehension10 0.6000.800 0.6000.800
Language-structure10 0.0000.100 0.0000.100
Language-transformation10 0.5000.600 0.5900.640
Logic-consistency5 0.2000.200 0.2000.200
Logic-deduction10 0.5000.500 0.5000.500
Logic-pattern10 0.4000.400 0.4000.400
Math-arithmetic8 0.6250.625 0.6250.625
Math-pattern10 0.8000.800 0.8000.800
Math-reasoning10 0.6000.500 0.6000.500
Pattern-generation9 0.7780.667 0.7780.667
Pattern-matching10 1.0000.800 1.0000.900
Pattern-recognition10 0.3000.300 0.3000.300

Artifacts

task: bench_mid_6_2026
dataset_path: json
dataset_name: null
dataset_kwargs:
  data_files:
    test: /home/user/eval.jsonl
output_type: multiple_choice
training_split: null
validation_split: null
test_split: test
num_fewshot: 0
doc_to_text: "Question: {{input}}\nAnswer:"
doc_to_choice: !function utils.doc_to_choice
doc_to_target: !function utils.doc_to_target
process_results: !function utils.process_results
metric_list:
  - metric: acc
    aggregation: mean
    higher_is_better: true
  - metric: acc_norm
    aggregation: mean
    higher_is_better: true
  - metric: soft_score
    aggregation: mean
    higher_is_better: true
  - metric: soft_score_norm
    aggregation: mean
    higher_is_better: true
metadata:
  version: 1.0

Evaluation run: lm-eval 0.4.12 · transformers 5.12.1 · device: CPU · batch size: 8 · samples: 155