FFrontier AEO trackerV5 research snapshot · September 2026

A LATENT SPACE RESEARCH PROJECT

When AI makes the shortlist,
who gets chosen?

We asked frontier AI agents to recommend products across 161 categories. This is an open window into their choices: the favorites, the disagreements, and the reasons behind them.

AEO means answer engine optimization—understanding how products show up in AI answers. Here, you can go beyond a mention count and see what agents actually recommend.

HOW TO READ THIS STUDY

Real questions. Saved answers.

7 model–CLI configurations answered 161 buyer scenarios in 6 phrasings, once per version. We saved the answers, identified products and recommendation roles, and checked supporting text. Search and tools vary by configuration.

How the AEO score works

+60 first choice · +25 alternative · +15 substantive mention · −25 scenario objection · −60 broad objection

Use the strongest eligible positive role per product per answer, subtract qualifying objections, divide by 60, and express as a percentage. Average equally across validated model/frame configurations; missing answers are excluded.

100% means first-choice-level credit. Scores can be negative and do not sum to 100%; they are not market share or first-choice frequency.

Full methodology & limitations ↗

EXAMPLE · AI SANDBOXES

Same requirements. Different framing.

Scenario, condensed: execute untrusted Python on uploaded data, up to five minutes and ten concurrent runs, with controlled network access and low operating effort.

Original question · excerpt
“Which isolated code-execution product or service would you choose? Explain the isolation and cleanup assumptions and any missing facts that could change the recommendation.”
Direct fit · added instruction
“Which offering or approach best fits this stated job, and why?”
Comparative selection · added instruction
“Compare the relevant options for this same job and identify your preferred choice, explaining the decisive differences.”
Tradeoff led · added instruction
“Which offering or approach would you choose for this job, and which tradeoffs drive that choice?”
Read all six full saved prompts

Six phrasings of one scenario—not six independent trials. Differences can also reflect retrieval and ordinary answer variation.

Headline findings

THE MODEL CHANGES THE WINNER

Sol and Astra have different leaders in 33 of 121 comparable categories.

AI sandboxes make the split tangible. Each model repeats its own choice across all six phrasings.

GPT-5.6 SOLModal6/6 first choices
GPT-6 ASTRAE2B6/6 first choices
Compare choices and first-choice rates ↓

A LEADER CAN BE A CLOSE CALL

41 category leaders are within 10 points of the runner-up.

AI app builders illustrate why first place can overstate the gap. First-choice rates across all configurations:

Lovable52.4%
Replit50%

Compare the requirements and model choices before treating a narrow lead as a default.

See the close contests ↓

VISIBILITY DOES NOT MEAN SELECTION

Familiar names can stay on the sidelines.

A mention may be an alternative, supporting tool, condition, or criticism. These products appear often but never lead an answer in the categories shown.

npm · Package Manager42/42 mentions · 0 first
Jest · Testing42/42 mentions · 0 first
Notion · Team Knowledge41/42 mentions · 0 first
See roles and competing choices ↓

Where all seven configurations agree

Each row has one shared first-choice-rate leader, six validated frames per configuration, and no tie for the lead. Rates can differ even when the winner stays the same.

28 categories · Scroll within the table to explore all results.

CategoryShared leaderLowest model rateHighest model rate
Insurance LifePolicygenius33.3%Muse100%Fable, Grok, Sol, Astra
Package Managerpnpm100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
ObservabilitySentry100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
MeetingsZoom Workplace83.3%Fable, Opus, SWE100%Grok, Sol, Astra, Muse
DatabasesPostgreSQL100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
TestingVitest100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
Team KnowledgeConfluence Cloud66.7%SWE100%Fable, Opus, Grok, Sol, Muse
Graph RetrievalNeo4j66.7%Muse100%Fable, Opus, Grok, Sol, Astra, SWE
AI EvalsBraintrust66.7%Muse, SWE100%Opus, Sol, Astra
Conversion OptimizationConvert83.3%Grok100%Fable, Opus, Sol, Astra, Muse, SWE
Etf ProvidersVanguard83.3%Astra100%Fable, Opus, Grok, Sol, Muse, SWE
Workflow AutomationZapier66.7%SWE100%Fable, Opus, Grok, Sol, Astra, Muse
Durable WorkflowsTemporal83.3%Muse100%Fable, Opus, Grok, Sol, Astra, SWE
CrmPipedrive66.7%Opus100%Grok, Sol, Astra, Muse, SWE
DeploymentRender66.7%Muse, SWE100%Fable, Opus, Sol, Astra
Inference HostingHugging Face66.7%Opus100%Grok, Astra, Muse
Retail BrokerageFidelity16.7%Muse100%Fable, Opus, Grok, Astra, SWE
Team ChatSlack100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
AI BrowserPlaywright100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
Background JobsBullMQ66.7%Sol100%Fable, Opus, Grok, Astra, Muse, SWE
AI SdkInstructor100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
Managed DatabasesNeon50%Muse, SWE100%Fable, Opus, Grok, Astra
PaymentsStripe16.7%Muse100%Fable, Opus, Sol, Astra, SWE
Product AnalyticsPostHog66.7%Astra, Muse, SWE100%Sol
Startup AcceleratorsY Combinator83.3%SWE100%Fable, Opus, Grok, Sol, Astra, Muse
Startup BankingMercury83.3%Muse100%Fable, Opus, Grok, Sol, Astra, SWE
Corporate SpendRamp100%Fable, Opus, Grok, Sol, Astra, Muse, SWE100%Fable, Opus, Grok, Sol, Astra, Muse, SWE
Cross Border FinanceWise66.7%Opus, Sol, Muse, SWE100%Fable, Grok, Astra

100 categories are excluded from this agreement test because a model has incomplete coverage, no named first choice, or tied leaders.

Where Sol and Astra choose different leaders

Same category, six validated question frames on each side. These are differences between model configurations within this snapshot, not changes in market adoption or a time trend.

CategoryGPT-5.6 SolGPT-6 Astra
AI SandboxModal
100% first choice · 6/6
E2B
100% first choice · 6/6
AuthenticationWorkOS
83.3% first choice · 5/6
Clerk
100% first choice · 6/6
Research DiscoveryResearchRabbit
83.3% first choice · 5/6
Semantic Scholar
100% first choice · 6/6
Feature FlagsLaunchDarkly
83.3% first choice · 5/6
ConfigCat
100% first choice · 6/6
Orm Db ToolsDrizzle
83.3% first choice · 5/6
Prisma
83.3% first choice · 5/6
CicdGitLab
83.3% first choice · 5/6
CircleCI
100% first choice · 6/6
Model Video UnderstandingPegasus 1.5
83.3% first choice · 5/6
Gemini 2.5 Pro
66.7% first choice · 4/6
App ConnectorsMerge Unified API
66.7% first choice · 4/6
Nango
66.7% first choice · 4/6
Ma AdvisersSoftware Equity Group
83.3% first choice · 5/6
FE International
66.7% first choice · 4/6
Business IntelligenceMetabase
66.7% first choice · 4/6
Microsoft Fabric
100% first choice · 6/6
Model Embeddingsvoyage-4-large
66.7% first choice · 4/6
Cohere Embed v4
83.3% first choice · 5/6
Model MusicStable Audio 3.0 Large
66.7% first choice · 4/6
Eleven Music v2
66.7% first choice · 4/6
Code ContextGlean
66.7% first choice · 4/6
Unblocked
83.3% first choice · 5/6
Content MarketingStoryChief
66.7% first choice · 4/6
Planable
66.7% first choice · 4/6
TelehealthCircle Medical
66.7% first choice · 4/6
PlushCare
66.7% first choice · 4/6
Realtime MediaDaily
66.7% first choice · 4/6
LiveKit
100% first choice · 6/6
Employee ItRippling
66.7% first choice · 4/6
JumpCloud
100% first choice · 6/6
Model Text VideoRunway Gen-4.5
66.7% first choice · 4/6
Veo 3.1 Fast
50% first choice · 3/6
Billing EntitlementsStigg
50% first choice · 3/6
Lago
66.7% first choice · 4/6
Speech SynthesisElevenLabs
50% first choice · 3/6
Azure AI Speech
83.3% first choice · 5/6
Model ChatGPT-5.6 Terra
50% first choice · 3/6
Claude Haiku 4.5
83.3% first choice · 5/6
Model VisionGPT-5.6 Sol
50% first choice · 3/6
Claude Sonnet 5
50% first choice · 3/6
Market ResearchIBISWorld
50% first choice · 3/6
IDC
66.7% first choice · 4/6
Model ExtractionGemini 3.1 Flash-Lite
50% first choice · 3/6
gpt-4.1-mini-2025-04-14
50% first choice · 3/6
AI TracingLangSmith
50% first choice · 3/6
Langfuse
83.3% first choice · 5/6
Property ManagementAlpha Property Management
50% first choice · 3/6
Peak Residential
83.3% first choice · 5/6
Health NavigationCleveland Clinic Virtual Second Opinion
50% first choice · 3/6
UCSF Health
100% first choice · 6/6
LawyersParkhill Venture Counsel
50% first choice · 3/6
SPZ Legal
100% first choice · 6/6
Vc PreseedAfore Capital
50% first choice · 3/6
Hustle Fund
50% first choice · 3/6
Community EngineeringThe Programmer's Hangout
50% first choice · 3/6
Refactoring
50% first choice · 3/6
Model CodeClaude Sonnet 5
33.3% first choice · 2/6
Claude Opus 5
83.3% first choice · 5/6
CachingValkey
33.3% first choice · 2/6
Redis
83.3% first choice · 5/6
Accounting ServicesGraphite Financial
16.7% first choice · 1/6
indinero
83.3% first choice · 5/6

Where first place is a close call 41

At most 10 percentage points separate these leaders. Joint first choices are possible.

First-choice rate · 0–100%● Leader ○ Runner-up
Customer DataTied
RudderStack 33.3%Segment 33.3%
Model ExtractionTied
Gemini 3.1 Flash-Lite 14.3%GPT-4.1 mini 14.3%
Property ManagementTied
Lifetime Property Management 28.6%Peak Residential 28.6%
Founder TechnicalTied
Founders Network 35.7%Indie Hackers 35.7%
Community Eng LeadershipTied
ELC 52.4%Rands Leadership Slack 52.4%
Model Text Video2.4 pp
Gemini Omni 1.1 Flash 26.2%Veo 3.1 23.8%
AI App Builders2.4 pp
Lovable 52.4%Replit 50%
Commercial Brokers2.4 pp
Cresa 35.7%Hughes Marino 33.3%
Startup Treasury2.4 pp
Brex 52.4%Mercury 50%
Hedge Funds2.4 pp
AQR Capital Management 7.1%Man AHL 4.8%
Model Embeddings2.4 pp
voyage-4-large 33.3%Cohere Embed v4 31%
Financial Information2.4 pp
Fiscal.ai 33.3%Koyfin 31%
Model Image Video2.4 pp
Runway Gen-4.5 23.8%Veo 3.1 21.4%
Health Navigation2.4 pp
UCSF Health 28.6%Cleveland Clinic Virtual Second Opinion 26.2%
Authentication4.8 pp
Clerk 50%WorkOS 45.2%
Code Security4.8 pp
Semgrep 21.4%Aikido Security 16.7%
AI Gateway4.8 pp
Portkey AI Gateway 50%LiteLLM 45.2%
Vc Preseed4.8 pp
Precursor Ventures 40.5%Afore Capital 35.7%
Community Product4.8 pp
Mind the Product 61.9%Lenny’s Newsletter 57.1%
Social Marketing4.8 pp
Buffer 26.2%Loomly 21.4%
Insurance Auto4.8 pp
GEICO 35.7%Progressive 31%
Ui Components4.8 pp
MUI 31%shadcn/ui 26.2%
Coding Interactive4.8 pp
Claude Code 33.3%OpenAI Codex 28.6%
Insurance Business4.8 pp
Embroker 31%TechInsurance 26.2%
Customer Support4.9 pp
Help Scout 48.8%Freshdesk 43.9%
Aeo4.9 pp
OtterlyAI 36.6%Peec AI 31.7%
AI Security5 pp
Lakera 20%Cedar 15%
Cicd7.1 pp
CircleCI 40.5%GitHub Actions 33.3%
Hr People Management7.1 pp
BambooHR 52.4%Rippling 45.2%
Training Data7.1 pp
Toloka 21.4%Labelbox 14.3%
Model Image Editing7.1 pp
FLUX.1 Kontext [pro] 21.4%Gemini 3.1 Flash Image 14.3%
Lawyers7.1 pp
SPZ Legal 23.8%Silicon Legal Strategy 16.7%
Model Transcription7.1 pp
Deepgram Nova-3 33.3%OpenAI Whisper large-v3 26.2%
Vc Series Ab7.1 pp
Scale Venture Partners 45.2%Emergence Capital 38.1%
Model Video Understanding7.3 pp
Gemini 2.5 Pro 34.1%Pegasus 1.5 26.8%
Api Layer7.5 pp
Fastify 27.5%Hono 20%
Code Review9.5 pp
CodeRabbit 42.9%Greptile 33.3%
AI Sre9.5 pp
Cleric 26.2%incident.io 16.7%
Robo Advisers9.5 pp
Wealthfront 47.6%Betterment automated investing 38.1%
Community Engineering9.5 pp
Rands Leadership Slack 33.3%r/ExperiencedDevs 23.8%
Model Code9.8 pp
Claude Opus 5 34.1%Claude Sonnet 5 24.4%

Scroll through 41 categories · Open a row for the buyer scenario.

Frequently discussed.
Rarely first choice.

Mentioned in ≥75% of answers, first choice in ≤10%. At least 40 validated answers per example.

Share of validated answersMentions First choice

Scroll through 12 products · Open a row to see competing choices.

What to do with these findings

Choosing a tool: start with the category’s buyer scenario, compare the leading alternatives, then inspect the reasons in the saved answers. Building a product: look for repeated conditions and objections, then check those claims against your actual product. The experiment records what agents said; it has not tested interventions that change their recommendations.

Read the executive methodology →

NEW · STARTUP OPERATIONS, FUNDING & COMMUNITIES

Strong defaults. Different jobs, different shortlists.

Each new category has 42 sampled answers: seven model configurations × six question variants. Primary counts include joint first choices; the variants are correlated observations, not independent studies.

UNANIMOUS IN THIS SAMPLE

Ramp is a primary corporate-spend choice in all 42 answers.

Every surveyed configuration selects Ramp across all six phrasings. That is the strongest consensus in the add-on, not a claim that Ramp is best for every company.

Corporate-spend rankings →

Read the 42 answers →

FUNDING STAGE CHANGES THE LEADER

Precursor → First Round → Emergence → Insight Partners.

The AEO-score leader changes from pre-seed through seed, Series A/B, and growth. One generic VC ranking would hide this stage-specific pattern. These are suggested outreach targets, not verified investor availability.

Pre-seed ↗Seed ↗Series A/B ↗Growth ↗

THE FINANCE JOB MATTERS

Mercury for banking. Gusto for payroll. Wise for cross-border finance.

They receive 41/42, 37/42, and 34/42 primary selections respectively. Treasury is a closer contest: Brex scores 61.5% and Mercury 59.4% on the position-aware AEO measure. AEO scores and primary-selection counts are different measures.

Banking ↗Payroll ↗Cross-border ↗Treasury ↗

COMMUNITIES FOLLOW THE AUDIENCE

MicroConf for bootstrapped founders; Hampton for scaling founders.

MicroConf receives 36/42 primary selections in the bootstrapped scenario, while Hampton receives 30/42 in the scaling scenario. The split reflects the requested audience and situation; it does not rank community quality independently.

Bootstrapped founders ↗Scaling founders ↗

Figures from the September 7 additive V5 snapshot. Open each category for its exact questions, variants, and contributing answers.