Kimi K2.6 is the only inference model in the room. LangGraph controls the workflow, not the prediction. A complete run executes 33 Winner, 10 Method, 20 Timing, and 5 Oversight agents in a strict conditional sequence, then publishes the result with receipts.
33 Winner10 Method20 Timing5 Oversight
Methodology v8 · engine facts current through 2026-07-23 · published probabilities remain AI-only, with market prices kept to benchmarking and diagnostics.
Direct Kimi K2.6 topology
The 33 + 10 + 20 + 5 Architecture
The number is not a marketing estimate. Production validation requires the exact topology before a forecast can be released. Each layer has one job, one conditional contract, and no authority to reach backward into an earlier probability.
33
Winner agents
P(winner)
Thirty-three independent Kimi K2.6 agents investigate seven weighted evidence categories. They alone carry winner-vote weight. Missing or inadmissible evidence becomes an explicit abstention, never an invented neutral vote.
10
Method agents
P(method | winner)
Five Kimi K2.6 method roles run once for Fighter A as winner and once for Fighter B as winner. The ten outputs price KO/TKO, submission, and decision while the winner point remains immutable.
20
Timing agents
P(round | winner, finish)
Five Kimi K2.6 timing roles run across A/B × KO/TKO/submission branches. The twenty outputs estimate survival-conditioned finish hazards for every scheduled round.
5
Oversight agents
audit · calibrate · hold
Five non-voting Kimi K2.6 agents audit evidence quality, correlation, historical calibration, disagreement, and uncertainty. They may lower confidence, widen the interval, or hold publication—but cannot rewrite probabilities.
68Kimi K2.6 agent executionsOne model. Four responsibilities. One coherent joint distribution.
How It Works
Every prediction follows the same seven-stage path. Each handoff narrows authority: evidence feeds Winner, Winner locks the marginal, Method and Timing fill conditional branches, and Oversight audits without voting.
01
Frozen Evidence Pack
Fight data is validated, normalized, freshness-gated, and frozen before inference. Incomplete fighters or failed evidence gates hold the run.
02
33 Winner Agents
Thirty-three Kimi K2.6 specialists answer independent evidence questions across seven fixed-weight lanes. Eligible reads reconcile into one market-independent winner probability.
03
Winner Point Locked
The winner marginal becomes immutable. No Method, Timing, Oversight, market, or presentation stage is allowed to reopen it.
04
10 Method Agents
Five Kimi K2.6 method roles run for each possible winner, producing conditional KO/TKO, submission, and decision distributions.
05
20 Timing Agents
Five Kimi K2.6 timing roles run across the four finish branches: Fighter A or B by KO/TKO or submission. They estimate round-by-round survival hazards.
06
5 Oversight Agents
Five non-voting Kimi K2.6 auditors examine evidence quality and uncertainty. Exact outcome cells are reconciled to 100% while preserving the locked winner marginal.
07
Public Ledger
A valid release is logged before the fight and measured after settlement with hit rate, Brier score, and engine-contract metadata.
What the Committee Weighs
Not every factor matters equally. The committee assigns each of its 33 voting specialist agents a weight based on how much that factor actually moves prediction accuracy, so a stylistic mismatch counts for more than a stat line. The 7 categories, weighted to 100%:
Camp environment, matchup-specific work, notice length, and relevant preparation partners are evaluated independently.
Published percentages are baseline importance shares. Evidence-grounding and validated worker-quality checks can reduce a weak lane's effective contribution before the complete committee is renormalized.
See all 33 factors
Stylistic / Matchup · 20%
Archetype clash (5%), Archetype clash ONLY: assess the historical interaction between the fighters' demonstrated styles, such as wrestler versus striker or pressure versus counter. Own the style-pair question, not current athleticism, camp, raw efficiency, or a full-fight win-condition synthesis.
Gameplan imposition (5%), Game-plan imposition ONLY: decide who can force the bout toward their preferred range, clinch, or ground phase and control its tempo. Do not re-score historical archetype priors, defensive neutralization, or every possible win condition.
Defensive interaction (5%), Defensive interaction ONLY: assess whether each fighter's demonstrated striking, takedown, submission, and positional defenses can neutralize the opponent's primary weapon. Do not grade general durability, offensive efficiency, or overall fight control.
Win condition (5%), Win conditions ONLY: identify the discrete, evidence-supported route or routes by which each fighter can realistically win and compare their accessibility. Do not create a generic factor summary or count an unsupported narrative as a path; use explicit matchup and finish-path evidence only.
Physical · 20%
Athleticism (4%), Athleticism ONLY: compare verified speed, explosiveness, and reaction-time evidence. Do not infer athleticism from age, record, reach, output, or archetype; if direct evidence is absent, abstain.
Cardio (4%), Cardio ONLY: assess sustainable pace, late-round output, and recovery over this bout's scheduled distance. Do not use activity level, age, or training-camp reputation as a substitute for cardio evidence.
Durability (4%), Durability ONLY: compare demonstrated chin, body durability, damage accumulation, and recovery after being hurt. Do not score offensive finishing ability, cardio, or injury speculation in this lane.
Size & leverage (4%), Size and leverage ONLY: assess verified reach, height, frame, and evidence-backed clinch or grappling leverage. Never invent fight-night mass or use an unofficial rehydration estimate; do not grade general technique.
Injury & biomechanics (4%), Injury and biomechanics ONLY: assess source-verified injuries, mobility limitations, movement restrictions, and accumulated physical wear. Do not infer an injury from a layoff, age, performance decline, or visual appearance; absent verified evidence means neutral.
Statistical · 20%
Striking efficiency (5%), Striking efficiency ONLY: compare accuracy, defense, landed-versus-absorbed differential, and verified knockdown production. Do not score pace alone, stylistic phase control, athleticism, or chin durability.
Grappling efficiency (4%), Grappling efficiency ONLY: compare takedown accuracy and defense, control production, submission activity, and verified scrambling efficiency. Do not decide the broader style clash or infer technical quality from discipline labels alone.
Pace & output (3%), Pace and output ONLY: compare normalized volume, tempo, and repeatable pressure rate over equivalent exposure. Do not grade accuracy, cardio reserve, or game-plan imposition in this lane.
Finishing metrics (3%), Finishing metrics ONLY: price KO/TKO and submission production separately against the opponent's finish resistance. Do not fold raw durability, stylistic access, or a generic method prediction into this statistical rate comparison.
Strength of competition (3%), Strength of competition ONLY: normalize records and observed performance for opponent quality using strength-of-schedule and rating evidence. Do not duplicate current efficiency metrics, career phase, or recent momentum.
Predictive metrics (2%), Predictive metrics ONLY: interpret the approved worker-visible deterministic statistical prior or composite derived metrics, including their sample and reliability limits. Never recreate a model from narrative, use persistence-only reconciliation fields, or borrow another lane when the approved prior is absent.
Film · 15%
Fight iq (3%), Fight IQ ONLY: judge in-cage decision making and tactical choices from source-attributed film observations. Results, record, method diversity, or a presumed game plan are not fight-IQ evidence; empty observations mean neutral.
Adaptability (3%), Adaptability ONLY: assess demonstrated mid-fight adjustments when an initial plan failed, using source-attributed film observations. Do not infer adaptability from career trajectory, comeback results, or multiple win methods; empty observations mean neutral.
Technical detail (3%), Technical detail ONLY: evaluate source-attributed mechanics, positioning, transitions, and defensive technique visible on tape. Do not substitute career statistics or broad style labels for a concrete technical observation.
Pattern recognition (3%), Pattern recognition ONLY: identify a recurring, source-supported tactical habit and whether this opponent can exploit it. A single outcome or one unsourced anecdote is not a recurring pattern; do not re-score general technique.
Hidden tape (3%), Hidden tape ONLY: assess source-attributed timing, composure, reads, and reactions that are not already represented by the statistical lanes. Do not invent intangibles or use popularity, narrative, record, or third-party picks; absent verified tape evidence means neutral.
Trajectory · 11%
Improvement (3%), Improvement ONLY: compare repeatable execution quality over recent, chronologically ordered fights. Paired completed-round-normalized last-three versus prior output/control trends are admissible development evidence; they do not prove a newly added technique. Do not treat a win streak, opponent downgrade, age, or one result as proof of improvement.
Decline (2%), Decline ONLY: assess evidence of sustained athletic regression or accumulated damage across normalized recent performances. Age alone, a single loss, or a long layoff is not decline evidence and belongs to other lanes.
Skill evolution (2%), Skill evolution ONLY: identify source-backed additions to the fighter's technical toolkit or strategic approach. Do not duplicate generic improvement, camp relationships, or presumed techniques inherited from a coach.
Psychological confidence (2%), Psychological confidence ONLY: assess source-supported momentum, composure, and confidence trends. Do not infer mental state from a win streak, crowd reaction, social media narrative, or the engine's prediction confidence; absent direct evidence means neutral.
Career arc (2%), Career arc ONLY: use the paired server-derived career_phase plus its observed age, professional-bout, and record inputs to compare prospect, prime, veteran, or championship-stage context. Do not duplicate recent improvement, decline, confidence, or strength-of-competition scoring.
Situational · 7%
Travel environment (1.5%), Travel and environment ONLY: assess verified travel distance, altitude, venue geography, climate, and time-zone adaptation evidence. Do not infer acclimatization or fatigue when the bundle supplies only a venue or home country.
Layoff activity (2%), Layoff and activity ONLY: compare inactivity duration, recent turnaround cadence, and the bounded risk of ring rust. Never assert why a fighter was inactive unless that reason is explicitly sourced, and do not convert inactivity into an injury or decline claim.
Weight cut (2%), Weight cut ONLY: assess prior official weigh-in history, missed-weight recency and severity, division changes, and actually measured size information. Do not use unofficial fight-night estimates or invent a current-event cut outcome.
Event context (1.5%), Event context ONLY: assess scheduled distance, verified main-event or championship-round experience, cage dimensions, and evidence-backed location context. Do not use crowd narrative, travel burden, cardio projection, or judging speculation as substitutes.
Preparation · 7%
Camp quality (2%), Camp quality ONLY: assess sourced gym, coaching staff, camp continuity, and verified disruptions. Matching source-verified active primary camps may support only a partial continuity equivalence, not a directional quality grade. Do not infer a clean or poor camp from missing data, and leave opponent-specific planning and partner fit to their own lanes.
Opponent preparation (2%), Opponent preparation ONLY: assess sourced evidence that the camp prepared specifically for this opponent's style or primary weapons. A coach relationship, generic camp statement, or plausible tactic is not matchup-specific preparation without an explicit sourced link.
Notice length (1.5%), Notice length ONLY: compare verified booking notice and full-camp versus short-notice status. Do not treat days since the last fight as notice, infer an opponent change, or grade camp quality.
Training partners (1.5%), Training partners ONLY: assess sourced partner identities and whether their demonstrated style supplies a relevant preparation look. Do not infer techniques from association, duplicate overall camp quality, or assume an unlisted partner was absent.
What the Committee Refuses to Weigh
A disciplined room is defined as much by what it ignores. These signals get zero weight: too unreliable, too situational, or too low-yield to move a serious forecast.
Zero weightWhat the committee refuses to weigh
staredown body language
unsupported motivation narratives
betting-market movement
unsourced referee and judging narratives
live/current-event scale weight
post-lock weigh-in reruns
unverified hydration speculation
raw intuition
Too unreliable, too situational, or too low-yield to move a serious forecast. No staredown ever changed who has the better chin.
The Engine's Week
The schedule is event-driven: establish a readiness-gated baseline, react only to durable material changes, run the promotion-aware final pass, and preserve the published forecast at the event's canonical lock.
01Early baseline
02Material delta if needed
03Final-ready pass
04Lock and backstop
After settlement01
Early baseline
Readiness-gated
After the prior card settles—or during the Sunday/Monday window—the system waits for a sufficiently complete, quiet card before freezing evidence and producing the first full read. A Tuesday safety valve prevents a ready card from being stranded.
After baseline02
Material-change watch
Every 30 minutes
The handoff monitor records roster and evidence changes durably. When a registered change is material, only the affected matchup is staged for a delta run; unchanged fights keep their frozen forecast.
Final-ready window03
Promotion-aware sensor
Every 15 minutes
During the configured Friday-to-event window, a promotion-aware sensor checks whether the card is ready for its final evidence pass. It self-deduplicates, detects stalled runs, and never treats current-event scale weight as prediction evidence.
Event lock04
Lock and preserve
Pre-doors + backstop
The event’s canonical lock time controls publication. A pre-doors scratch sweep catches late removals, while a Saturday/Sunday UTC fallback protects against a missed sensor window. Locked forecasts are preserved for grading.
The baseline monitor runs every 30 minutes; the configured final-ready windows run every 15 minutes. Exact timing follows the promotion, card state, and canonical event lock. After lock, the published forecast is preserved for grading.
What Makes This Different
Single-pass models compress an entire fight into one answer. The Oracle separates the questions. 33 Winner agents establish direction, 10 Method agents price how each fighter could win, 20 Timing agents price when a finish could arrive, and 5 non-voting agents audit whether the evidence is strong enough to publish.
Long-context committee reasoning keeps the fight packet, source context, and tool evidence in view
Disagreement is promoted to the reconciler instead of smoothed away
The result is not more noise. It is more independent scrutiny before one number is allowed to go public.
No Black Box Claims
You are not asked to “trust the model.” You can verify it.
Predictions are published before fights
Engine metadata is logged with each release
Hit rate and Brier score are measured after settlement
Accuracy and calibration are not claims. They are recorded outcomes.
Built For Real Decisions
This is not a content feed.
It is a system designed to produce consistent forecasts, maintain probability discipline, and operate under real-world conditions without blending sportsbook or exchange prices into the engine read.
Every output is constrained, calibrated, and logged.
The Standard
If a prediction cannot be validated, it does not belong here.
Data integrity requirements
Live-signal staleness gates
Kimi K2.6 specialist disagreement checks, deterministic reconciliation, and calibration against settled outcomes
Anything that fails is not surfaced.
Continuously Improved
The system is actively audited to:
Improve specialist coverage without duplicating viewpoints
Keep probability calibration honest against settled fights
Preserve public proof while keeping internal prompts, routing, and thresholds private
Changes are kept only if they improve validated results.
Accountable Predictions
This is not about picking winners. It is about producing accountable predictions.
Every forecast is data-driven
Every forecast is validated before release
Every forecast is logged with engine metadata and measured after the fact
Public Track Record: every settled prediction with hit rate and Brier score.
FAQ
What happens before a forecast is published?
Fight data is validated and frozen under freshness gates, then the direct Kimi K2.6 pipeline runs 68 Kimi K2.6 agent executions: 33 Winner, 10 Method, 20 Timing, and 5 non-voting Oversight. Winner probability is fixed first; Method and Timing agents condition on that result; Oversight can lower confidence or hold publication but cannot rewrite the probabilities. The coherent joint distribution is held to the 5-95% winner bound and logged before the bell. Market odds are comparison data only. The independent33 roster may use source-verified historical weight-cut trends and travel/environment context available before the prediction lock. It does not ingest current-event scale weight, hydration speculation, or trigger live/post-lock weigh-in reruns.
How does the committee decide what matters most?
The seven category weights total 100%: stylistic 20%, physical 20%, statistical 20%, film 15%, trajectory 11%, situational 7%, and preparation 7%. Unsupported, stale, invalid, or absent domain evidence produces a null abstention. Eligible specialists renormalize only within the same lane when at least two specialists and 50% nominal coverage remain; unresolved mass lowers confidence and widens the interval. The five oversight agents remain at zero vote weight. The independent33 roster may use source-verified historical weight-cut trends and travel/environment context available before the prediction lock. It does not ingest current-event scale weight, hydration speculation, or trigger live/post-lock weigh-in reruns.
Why does Blueprint call this a 68-agent architecture?
A complete topology contains 33 Winner agent executions, 10 Method agent executions, 20 Timing agent executions, and 5 Oversight agent executions: 68 total. The Method and Timing layers reuse five tightly defined role identities across conditional branches, but each branch is an independent Kimi K2.6 execution with its own evidence projection, output contract, validation, and persisted trace.
Is Kimi K2.6 the only AI model making the prediction?
Yes. Kimi K2.6 is the only inference model in this prediction architecture. LangGraph coordinates stage order, state, retries, and failure handling; it is orchestration software, not another forecasting model. The direct pipeline fails closed instead of switching to a legacy or substitute prediction model.
Do subscription tiers change model behavior?
No. Every tier reads the same Kimi K2.6 architecture. Plans only affect access to full predictions and workflow tools.
How is performance verified?
Predictions are logged before fights and compared against outcomes after settlement. Hit rate, Brier score, confidence behavior, and engine metadata are tracked on the public Track Record.
What happens if the engine cannot produce a valid read?
The prediction is held. The direct Kimi K2.6 pipeline fails closed instead of substituting another model, a legacy prediction engine, or a stale artifact.