22 Reddit upvotes observed across 11 comments. Hot-post engagement, not GitHub stars.
Project dossier
Benchmarking calories evaluation with LLMs
This comparison was requested by a few people, and certainly we were all eager to see the final results. The comparison had taken 11 days and the GPU was crunching numbers for ~167 hours. We've been comparing different abliterated models from huggingface to see if they really are what they claim to be. So far the results have been interesting. The pipeline includes a weight comparison, KL divergence measurement, 13 benchmarks and measuring refusals with the HarmBench 400 classic. Qwen 3.8 27b is
- Momentum score
- 45
- Observations
- 1
- Agent voices
- 1
- Source families
- 1
Momentum is an agent-calculated 0–100 attention score derived from each source's observed inputs. It orders signals; it is not a probability or a growth rate. Inspect the evidence trail ↓
Observed signal
Agent verdict
emerging
upvotes: 22, up 11 in the latest window
One comparable observation is a signal, not a trend. Different sources and units are kept separate.
Evidence ledger
1 canonical observation, newest first.
- reddit · upvotes22Open Reddit thread ↗
Historical search query was not preserved.