Headline findings
SHARED DEFAULTS
28 categories have the same leader across all seven configurations.
That is 28 of 61 categories with complete coverage and one clear leader per model. Agreement identifies choices that survive a change of model; it does not establish which product is best.
pnpm · Package ManagerSentry · ObservabilityZoom Workplace · MeetingsPostgreSQL · DatabasesVitest · TestingConvert · Conversion Optimization
See the agreement table ↓THE MODEL CHANGES THE WINNER
Sol and Astra have different leaders in 33 of 121 comparable categories.
AI sandboxes make the split tangible. Each model repeats its own choice across all six phrasings.
A LEADER CAN BE A CLOSE CALL
41 category leaders are within 10 points of the runner-up.
AI app builders illustrate why first place can overstate the gap. First-choice rates across all configurations:
Compare the requirements and model choices before treating a narrow lead as a default.
See the close contests ↓VISIBILITY DOES NOT MEAN SELECTION
Familiar names can stay on the sidelines.
A mention may be an alternative, supporting tool, condition, or criticism. These products appear often but never lead an answer in the categories shown.
Where all seven configurations agree
Each row has one shared first-choice-rate leader, six validated frames per configuration, and no tie for the lead. Rates can differ even when the winner stays the same.
28 categories · Scroll within the table to explore all results.
| Category | Shared leader | Lowest model rate | Highest model rate |
|---|---|---|---|
| Insurance Life | Policygenius | 33.3%Muse | 100%Fable, Grok, Sol, Astra |
| Package Manager | pnpm | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Observability | Sentry | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Meetings | Zoom Workplace | 83.3%Fable, Opus, SWE | 100%Grok, Sol, Astra, Muse |
| Databases | PostgreSQL | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Testing | Vitest | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Team Knowledge | Confluence Cloud | 66.7%SWE | 100%Fable, Opus, Grok, Sol, Muse |
| Graph Retrieval | Neo4j | 66.7%Muse | 100%Fable, Opus, Grok, Sol, Astra, SWE |
| AI Evals | Braintrust | 66.7%Muse, SWE | 100%Opus, Sol, Astra |
| Conversion Optimization | Convert | 83.3%Grok | 100%Fable, Opus, Sol, Astra, Muse, SWE |
| Etf Providers | Vanguard | 83.3%Astra | 100%Fable, Opus, Grok, Sol, Muse, SWE |
| Workflow Automation | Zapier | 66.7%SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse |
| Durable Workflows | Temporal | 83.3%Muse | 100%Fable, Opus, Grok, Sol, Astra, SWE |
| Crm | Pipedrive | 66.7%Opus | 100%Grok, Sol, Astra, Muse, SWE |
| Deployment | Render | 66.7%Muse, SWE | 100%Fable, Opus, Sol, Astra |
| Inference Hosting | Hugging Face | 66.7%Opus | 100%Grok, Astra, Muse |
| Retail Brokerage | Fidelity | 16.7%Muse | 100%Fable, Opus, Grok, Astra, SWE |
| Team Chat | Slack | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| AI Browser | Playwright | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Background Jobs | BullMQ | 66.7%Sol | 100%Fable, Opus, Grok, Astra, Muse, SWE |
| AI Sdk | Instructor | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Managed Databases | Neon | 50%Muse, SWE | 100%Fable, Opus, Grok, Astra |
| Payments | Stripe | 16.7%Muse | 100%Fable, Opus, Sol, Astra, SWE |
| Product Analytics | PostHog | 66.7%Astra, Muse, SWE | 100%Sol |
| Startup Accelerators | Y Combinator | 83.3%SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse |
| Startup Banking | Mercury | 83.3%Muse | 100%Fable, Opus, Grok, Sol, Astra, SWE |
| Corporate Spend | Ramp | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE | 100%Fable, Opus, Grok, Sol, Astra, Muse, SWE |
| Cross Border Finance | Wise | 66.7%Opus, Sol, Muse, SWE | 100%Fable, Grok, Astra |
100 categories are excluded from this agreement test because a model has incomplete coverage, no named first choice, or tied leaders.
Special interest · Two deep dives
What makes a model choose you?
Read the recommendations, the cited evidence, and what website owners can learn from both.
Being known doesn’t mean you qualify.
Prove the capability that could rule you out. The sandbox answers show how a missing fact can cost a product consideration—and why proving the feature still may not win the recommendation.
Capability pages · Implementation guides · Show what your evidence provesRead Sol / Astra Claude Opus 5 / Claude Fable 5.1Give the specialist advantage a reason to matter.
Show when your strength is worth the tradeoff. Filtering, sync, and tracing reveal when specialists win, when platforms win, and what persuades both models.
Specialist advantages · Shared winners · Benchmark limitsRead Opus / FableWhere Sol and Astra choose different leaders
Same category, six validated question frames on each side. These are differences between model configurations within this snapshot, not changes in market adoption or a time trend.
| Category | GPT-5.6 Sol | GPT-6 Astra |
|---|---|---|
| AI Sandbox | Modal 100% first choice · 6/6 | E2B 100% first choice · 6/6 |
| Authentication | WorkOS 83.3% first choice · 5/6 | Clerk 100% first choice · 6/6 |
| Research Discovery | ResearchRabbit 83.3% first choice · 5/6 | Semantic Scholar 100% first choice · 6/6 |
| Feature Flags | LaunchDarkly 83.3% first choice · 5/6 | ConfigCat 100% first choice · 6/6 |
| Orm Db Tools | Drizzle 83.3% first choice · 5/6 | Prisma 83.3% first choice · 5/6 |
| Cicd | GitLab 83.3% first choice · 5/6 | CircleCI 100% first choice · 6/6 |
| Model Video Understanding | Pegasus 1.5 83.3% first choice · 5/6 | Gemini 2.5 Pro 66.7% first choice · 4/6 |
| App Connectors | Merge Unified API 66.7% first choice · 4/6 | Nango 66.7% first choice · 4/6 |
| Ma Advisers | Software Equity Group 83.3% first choice · 5/6 | FE International 66.7% first choice · 4/6 |
| Business Intelligence | Metabase 66.7% first choice · 4/6 | Microsoft Fabric 100% first choice · 6/6 |
| Model Embeddings | voyage-4-large 66.7% first choice · 4/6 | Cohere Embed v4 83.3% first choice · 5/6 |
| Model Music | Stable Audio 3.0 Large 66.7% first choice · 4/6 | Eleven Music v2 66.7% first choice · 4/6 |
| Code Context | Glean 66.7% first choice · 4/6 | Unblocked 83.3% first choice · 5/6 |
| Content Marketing | StoryChief 66.7% first choice · 4/6 | Planable 66.7% first choice · 4/6 |
| Telehealth | Circle Medical 66.7% first choice · 4/6 | PlushCare 66.7% first choice · 4/6 |
| Realtime Media | Daily 66.7% first choice · 4/6 | LiveKit 100% first choice · 6/6 |
| Employee It | Rippling 66.7% first choice · 4/6 | JumpCloud 100% first choice · 6/6 |
| Model Text Video | Runway Gen-4.5 66.7% first choice · 4/6 | Veo 3.1 Fast 50% first choice · 3/6 |
| Billing Entitlements | Stigg 50% first choice · 3/6 | Lago 66.7% first choice · 4/6 |
| Speech Synthesis | ElevenLabs 50% first choice · 3/6 | Azure AI Speech 83.3% first choice · 5/6 |
| Model Chat | GPT-5.6 Terra 50% first choice · 3/6 | Claude Haiku 4.5 83.3% first choice · 5/6 |
| Model Vision | GPT-5.6 Sol 50% first choice · 3/6 | Claude Sonnet 5 50% first choice · 3/6 |
| Market Research | IBISWorld 50% first choice · 3/6 | IDC 66.7% first choice · 4/6 |
| Model Extraction | Gemini 3.1 Flash-Lite 50% first choice · 3/6 | gpt-4.1-mini-2025-04-14 50% first choice · 3/6 |
| AI Tracing | LangSmith 50% first choice · 3/6 | Langfuse 83.3% first choice · 5/6 |
| Property Management | Alpha Property Management 50% first choice · 3/6 | Peak Residential 83.3% first choice · 5/6 |
| Health Navigation | Cleveland Clinic Virtual Second Opinion 50% first choice · 3/6 | UCSF Health 100% first choice · 6/6 |
| Lawyers | Parkhill Venture Counsel 50% first choice · 3/6 | SPZ Legal 100% first choice · 6/6 |
| Vc Preseed | Afore Capital 50% first choice · 3/6 | Hustle Fund 50% first choice · 3/6 |
| Community Engineering | The Programmer's Hangout 50% first choice · 3/6 | Refactoring 50% first choice · 3/6 |
| Model Code | Claude Sonnet 5 33.3% first choice · 2/6 | Claude Opus 5 83.3% first choice · 5/6 |
| Caching | Valkey 33.3% first choice · 2/6 | Redis 83.3% first choice · 5/6 |
| Accounting Services | Graphite Financial 16.7% first choice · 1/6 | indinero 83.3% first choice · 5/6 |
Where first place is a close call 41
At most 10 percentage points separate these leaders. Joint first choices are possible.
Scroll through 41 categories · Open a row for the buyer scenario.
Frequently discussed.
Rarely first choice.
Mentioned in ≥75% of answers, first choice in ≤10%. At least 40 validated answers per example.
Scroll through 12 products · Open a row to see competing choices.
NEW · STARTUP OPERATIONS, FUNDING & COMMUNITIES
Strong defaults. Different jobs, different shortlists.
Each new category has 42 sampled answers: seven model configurations × six question variants. Primary counts include joint first choices; the variants are correlated observations, not independent studies.
UNANIMOUS IN THIS SAMPLE
Ramp is a primary corporate-spend choice in all 42 answers.
Every surveyed configuration selects Ramp across all six phrasings. That is the strongest consensus in the add-on, not a claim that Ramp is best for every company.
Corporate-spend rankings →FUNDING STAGE CHANGES THE LEADER
Precursor → First Round → Emergence → Insight Partners.
The AEO-score leader changes from pre-seed through seed, Series A/B, and growth. One generic VC ranking would hide this stage-specific pattern. These are suggested outreach targets, not verified investor availability.
THE FINANCE JOB MATTERS
Mercury for banking. Gusto for payroll. Wise for cross-border finance.
They receive 41/42, 37/42, and 34/42 primary selections respectively. Treasury is a closer contest: Brex scores 61.5% and Mercury 59.4% on the position-aware AEO measure. AEO scores and primary-selection counts are different measures.
COMMUNITIES FOLLOW THE AUDIENCE
MicroConf for bootstrapped founders; Hampton for scaling founders.
MicroConf receives 36/42 primary selections in the bootstrapped scenario, while Hampton receives 30/42 in the scaling scenario. The split reflects the requested audience and situation; it does not rank community quality independently.
Figures from the September 7 additive V5 snapshot. Open each category for its exact questions, variants, and contributing answers.