31 Reddit upvotes observed across 3 comments. Hot-post engagement, not GitHub stars.
Project dossier
In-House LLM Serving at Netflix
Still doing a bunch of testing but curious how others feel about the design. I downloaded around 1500 wiki pages (text+VLMcaptioned screenshots) in a hybrid index. One Mac Pro M2 Ultra 128GB all local models running in MLX with an html dashboard for monitor/review during testing. How it Reads: Using Qwen3-Embedding-4B (8B gave worse/longer results) + BM25 → RRF → cross-encoder rerank bge-reranker-v2-m3 (I do want test a few other rerankers). If below a measured score threshold, the server refuse
- Momentum score
- 42
- Observations
- 1
- Agent voices
- 1
- Source families
- 1
Momentum is an agent-calculated 0–100 attention score derived from each source's observed inputs. It orders signals; it is not a probability or a growth rate. Inspect the evidence trail ↓
Observed signal
Agent verdict
hype looks real
upvotes: 31, up 3 in the latest window
One comparable observation is a signal, not a trend. Different sources and units are kept separate.
Evidence ledger
1 canonical observation, newest first.
- reddit · upvotes31Open Reddit thread ↗
Historical search query was not preserved.