Skip to content

Payments Intelligence

System Metrics

How healthy and trustworthy the intelligence system is — not what the payments market is doing (that is Market Pulse). Every figure counts only publishable briefs: confidence ≥ 40, three or more ranked sources, no “insufficient evidence” refusal.

64.2
Average confidence
63.9
Average impact
85 publishable briefs · all time121 companies tracked18 at impact 70+

Primary-source share 18.3% · full article text retrieved 99.8%. Scores are model-assigned 0–100 — 40–59 Moderate, 60–79 High, 80+ Very high.

Payments Intelligence

By the numbers

Measured since 28 Aug 2026 · all time shown

479
Source observations
176
Autonomous requests
71
Published briefs
97.7%
Autonomous completion
176 of 182 runs
636
Offline evaluations
Not measured yet
Analyst hours automated
Awaiting a timed analyst benchmark

Every figure above is either directly counted (OBSERVED), arithmetic over counted facts (DERIVED), or derived from a recorded benchmark (ESTIMATED, never fabricated). A request the primary research engine hands to the resilient fallback is one autonomous request, not two — see the methodology below.

How these numbers are calculated
Source observationsOBSERVED

One row per source cited in one published or duplicate-resolved brief. The same URL cited across 5 research runs is 5 observations, not 1.

Exclusions: Withheld and failed runs never produce cited sources, so they add nothing here.

Not the same as unique sources — see unique source URLs.

Source: public.brief_sources, public.pipeline_runs

Autonomous requestsOBSERVED

Every research request the system handled, counted once. When the primary research engine hands a request to the resilient fallback, both engine runs share one request id, so the pair is exactly ONE request, never two. Requests answered directly by the fallback (for example when the daily request budget was reached before the primary engine could attempt them) are real requests and are counted once as well.

Exclusions: Controlled production evaluation runs (scheduled and manual evaluation of the primary engine) are excluded by construction — they are evaluation traffic, not requests.

An earlier version of this formula subtracted every fallback link, including requests the primary engine never attempted, and so under-counted real requests — corrected when the fallback-rate review surfaced it.

Source: public.pipeline_runs, public.pipeline_fallback_links

Published briefsOBSERVED

A real, successful publication only. Duplicate, withheld and failed outcomes are separate, named categories, never folded into this count.

Exclusions: Duplicates, withholds and failures are shown as their own figures.

Source: public.pipeline_runs

Autonomous completionDERIVED

Of the engine runs whose origin is KNOWN to be automatic (authenticated ingress, on-demand form, daily monitor or schedule), the share that completed without a person stepping in. A run that had to be swept as stalled counts as not completed; that sweep signal is never used to decide which runs count as known, and is never treated as a proxy for manual intervention.

Exclusions: Runs whose origin is manual or unknown are left out entirely — they cannot honestly be assumed automatic or manual either way.

No known-origin runs in range -> shown as "Not measured yet", never a fabricated percentage. See automation coverage for how much of all runs this rate speaks to.

Source: public.pipeline_runs, public.impact_metrics_daily

Automation coverageDERIVED

The share of all engine runs whose execution origin is known at all — how much of the system autonomous completion actually describes.

Low coverage means autonomous completion describes a small, possibly unrepresentative slice — which is why it is shown alongside the rate.

Source: public.pipeline_runs, public.impact_metrics_daily

Offline evaluationsOBSERVED

The real replay volume: every {candidate, historical observation} comparison the offline evaluator actually executed, summed across cycles. For example 3 candidates x 400 historical observations actually replayed = 1,200 — but only because 1,200 comparisons really ran; it is never candidate_count * historical_count computed after the fact.

Exclusions: Counts cycles since this telemetry started only; earlier replay volume is not recoverable and is not estimated.

"Candidates on file" on the Learning page is a separate, smaller figure, not the same thing.

Source: public.learning_cycles

Eligible research requestsOBSERVED

The population Analyst hours automated is computed over: autonomous research requests that actually performed comparable research — the topic was researched, sources were ranked and read, and a publish-or-withhold decision was made. Published, duplicate and withheld outcomes all count (each is real research); a request the primary engine handed to the fallback counts once.

Exclusions: Failed or stalled runs, runs of manual or unknown origin (including tests), controlled evaluation runs and shadow evaluations never count.

Conservative by design: a fallback run whose primary attempt has an unknown origin is excluded rather than guessed.

Source: public.pipeline_runs, public.pipeline_fallback_links

Analyst hours automatedESTIMATED

Eligible research requests x the median minutes a human analyst needed for the same end-to-end research task, / 60. The median comes from a recorded benchmark of timed analyst tasks; it applies to every eligible request in the window, including those before the benchmark was recorded.

Always ESTIMATED. Shows "Not measured yet" until a real benchmark is recorded — never a guessed baseline or a fabricated 0. The benchmark type (measured or assumption), sample size and date are shown with the figure.

Source: public.research_benchmarks, public.pipeline_runs, public.pipeline_fallback_links

System scale & autonomy

What the autonomous system has actually done

Deduplicated, source-attributed figures for all time — never a raw internal workflow-execution count.

Scale
  • Source observations479
  • Unique source URLs210
  • Unique domains118
  • Primary-source share16.1%
Autonomy & reliability
  • Primary engine attempts3
  • Controlled evaluation runs (excluded)4
  • Fallback invocations5
  • Fallback rate66.7%(2 of 3)
  • Publish success rate55.5%
  • Automation coverage96.7%(176/182)
Learning & evaluation
  • Learning experiences21
  • Source reputation domains13
  • Candidates on file6
  • Passed offline1
  • Failed offline5
46s median research time62s p95 research time167 timed runs sampled

Candidates that FAILED_OFFLINE are the safety gate working correctly, not a bug — see the Learning page for the full candidate list and source-quality detail.

Publication quality

How well-evidenced the published stream is

Score profile and source quality for briefs that cleared the gate in all time.

Confidence distribution
  • 0–200
  • 20–400
  • 40–6027
  • 60–8053
  • 80–1005

Only briefs that clear the gate are stored, so the low bands are expected to be empty.

Impact distribution
  • 0–200
  • 20–403
  • 40–6010
  • 60–8072
  • 80–1000

18 briefs at impact 70+ · average 63.9.

Source-tier composition
  • Tier 1 — regulators / central banks13.3% · 53
  • Tier 2 — company primary5% · 20
  • Tier 3 — trade & business press30.5% · 122
  • Tier 4 — other51.3% · 205
Recent publishing
DayBriefsAvg confidenceAvg impactTier-1 sources
27 Sept 2026357.3661
26 Sept 2026268.5585
25 Sept 202636565.34
24 Sept 2026261.5650
23 Sept 2026268.558.50
22 Sept 2026372624
21 Sept 2026165480
20 Sept 2026258.5650
19 Sept 2026371.7640
18 Sept 2026165680

Coverage footprint

What the stream is made of

The category mix and most-cited domains. For rising themes and active entities, see Market Pulse.

Categories
  • Partnership29
  • Product23
  • Product14
  • Regulation10
  • M&A9
Most-cited sources
  • ecb.europa.eut1 · 50
  • fintechmagazine.comt3 · 31
  • cnbc.comt3 · 23
  • nuvei.comt4 · 19
  • thefintechtimes.comt3 · 15
  • reuters.comt3 · 13
  • adyen.comt2 · 12
  • thenextweb.comt4 · 12

System operations

Publication-gate outcomes and pipeline health

A different table from the quality figures above — the pipeline-run log, read from a separate log of runs. Operational health, not a market metric.

190 pipeline runs74 published40.7% publication rate4.2% failure share
OutcomeRunsShare of decided
Published7440.7%
Duplicate (story already published)3418.7%
Withheld — refused6636.3%
Withheld — insufficient evidence73.8%
Withheld — low confidence10.5%
Failed84.4%

A withheld run (model refusal, thin evidence, low confidence) is the system working correctly, not a failure. 182 decided runs · median run 46s · 5.6 sources retrieved per source used. One publication policy applies to every run. No per-run detail is exposed here.

today latest brief26 Aug 2026 first brief85 publishable briefs, all-time5 weekly reports

Latest publishable brief: 27 Sept 2026, 06:42. Most recent weekly report: Week 39, 2026. Method and limitations are on the Methodology page.

Engine comparison

Primary research engine and resilient fallback

Which research engine produced each run, all time. Authenticated requests go to the primary research engine first; the resilient fallback engine is the long-established pipeline that answers whenever the primary engine withholds or cannot run.

EngineRunsPublishedWithheldPublication rateAvg confidence
Resilient fallback engine181677466.3%68
Primary research engine970100%61.4

Each engine's own withheld runs (refused, insufficient evidence, low confidence) are the system working correctly, not a failure. No per-run or per-topic detail is exposed here.

Primary engine & resilient fallback

How often the primary engine needed the fallback, and how that went

Every request handled by the primary research engine, all time — whether it resolved on its own or was handed to the resilient fallback, paired by a request id assigned at ingress, never a topic/time guess.

8
Primary executions
2
Resolved without fallback
75%
Fallback rate
6 of 8 executions
16.7%
Fallback success
Primary vs fallback outcome
  • Not classified6
  • Resolved without fallback2

A factual outcome relationship, never a “which engine is better” score.

60 avg confidence, resolved without fallback2 primary engine attempted before falling back4 answered directly by the fallback (primary not attempted: gate closed or budget reached)65 avg confidence of fallback answersLast fallback: 26 Sept 2026, 18:42.