How often did AI models disagree in our 75-question test?
In this exploratory sample, the available model outputs produced materially different recommendations on 30 of 75 questions. That is 40% of this test set—not 40% of all AI questions.
Opposing bottom lines require evidence or qualified human review.
Disagreement by domain
“Disagreement” combines Partial results—same broad direction with a caveat that could change an action—and Split results with opposing bottom-line recommendations.
Domain
Questions
Partial + split
Rate in this sample
Life decisions
17
10
59%
Health
16
8
50%
Legal
13
6
46%
Finance
13
3
23%
Predictions
10
2
20%
Consumer choices
6
1
17%
Where AI models disagree most
Disagreement rises when the decision matters most.
Share of questions classified as Partial or Split. The counts come directly from the 75-question JSON dataset; they are not estimates of all AI questions.
Life decisions
10 of 17 questions
59%
Health
8 of 16 questions
50%
Legal
6 of 13 questions
46%
Finance
3 of 13 questions
23%
Predictions
2 of 10 questions
20%
Consumer choices
1 of 6 questions
17%
0% disagreement60% disagreement
Life decisions were 3.5× more likely to diverge than consumer choices: 59% versus 17% in this exploratory sample.
Method
1. Question set. 75 curated questions across health, law, finance, life decisions, predictions and consumer choices. The sample was not random.
2. Common answer prompt. Each attempted provider received: “You are a knowledgeable assistant. Answer the user's real-life question directly and give a clear bottom-line recommendation. Be concise (max 130 words).”
3. Attempted panel.Claude (claude-sonnet-4-6); ChatGPT (gpt-4o); Gemini (gemini-2.5-flash); Mistral (mistral-large-latest); Perplexity (sonar); Grok (grok-3 (attempted, not represented as a successful sixth response in the scored rows)).
4. Classification. One Claude Sonnet 4.6 judge classified the available outputs as Agree, Partial or Split and assigned a 0–100 alignment score.
All 75 classified questions
These are historical study outputs, not current recommendations. Open a row to inspect the recorded stance summaries.
1. health
Is it safe to take ibuprofen if I'm taking lisinopril for high blood pressure?
PARTIAL72/100
Recorded divergence: ChatGPT says occasional use is 'typically fine' while Claude, Gemini, and Perplexity frame it as generally unsafe or not recommended, creating a materially different default stance on casual use.
Claude
no - avoid, use acetaminophen
ChatGPT
depends - occasional use fine
Gemini
no - not recommended, ask doctor
Mistral
depends - short-term sometimes acceptable
Perplexity
no - generally not safe, use acetaminophen
Grok
N/A
5 successful provider outputs recorded for this question.
2. health
I've had a sore throat for 4 days with a mild fever. Should I take antibiotics?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
no - get tested first
ChatGPT
no - see doctor first
Gemini
no response provided
Mistral
no - monitor then test
Perplexity
no - get tested first
Grok
no response provided
4 successful provider outputs recorded for this question.
3. health
Is intermittent fasting safe and effective for a 55-year-old with type 2 diabetes?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
yes, with medical supervision
ChatGPT
yes, with medical supervision
Gemini
yes, with medical supervision
Mistral
yes, with medical supervision
Perplexity
yes, with medical supervision
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
4. health
Should a healthy 30-year-old get the annual flu vaccine?
AGREE99/100
No material bottom-line divergence was recorded by the judge.
Claude
yes – safe widely recommended
ChatGPT
yes – protects self and community
Gemini
yes – reduces spread protects vulnerable
Mistral
yes – CDC recommends for all
Perplexity
yes – high efficacy low exceptions
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
5. legal
In France, can my landlord keep my security deposit for normal wear and tear after I move out?
AGREE98/100
No material bottom-line divergence was recorded by the judge.
Claude
no, cannot keep deposit
ChatGPT
no, cannot keep deposit
Gemini
no, cannot keep deposit
Mistral
not present
Perplexity
no, cannot keep deposit
Grok
not present
4 successful provider outputs recorded for this question.
6. legal
If a guest slips on the icy walkway outside my house, am I legally liable?
PARTIAL72/100
Recorded divergence: Perplexity leads with 'not automatically liable' and emphasizes natural accumulation protection, while Gemini and Mistral lean more affirmatively toward likely liability, creating materially different risk framing for the homeowner.
Claude
depends + clear salt document insure
ChatGPT
depends + clear salt walkways
Gemini
yes + clear salt warn guests
Mistral
yes + clear salt warn document
Perplexity
not automatically + depends on cause/timing
Grok
N/A
5 successful provider outputs recorded for this question.
7. legal
Can my employer in France legally require me to work on a public holiday?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
yes, except May 1st, check convention collective
ChatGPT
depends, check contract and collective agreement
Gemini
yes, check contract and collective agreement
Mistral
yes except May 1st, check contract or agreement
Perplexity
yes for non-May-1st holidays, check contract/agreement
Grok
N/A
5 successful provider outputs recorded for this question.
8. finance
I have a 4% mortgage and €30k savings. Should I pay down the mortgage early or invest in an index fund?
AGREE82/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - invest, split suggested
ChatGPT
depends - invest if risk-tolerant
Gemini
yes - invest in index fund
Mistral
depends - split or invest
Perplexity
yes - invest in index fund
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
9. finance
I drive 15,000 km a year. Is it financially smarter to lease or buy a new car?
AGREE85/100
No material bottom-line divergence was recorded by the judge.
Claude
buy – unambiguous recommendation
ChatGPT
buy – with mild hedging
Gemini
buy – clear recommendation
Mistral
buy – with lease caveat for lifestyle preference
Perplexity
buy – despite confused km/miles analysis
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
10. finance
Should I withdraw from my retirement account to pay off €15k of credit-card debt at 20% APR?
SPLIT35/100
Recorded divergence: Gemini explicitly recommends withdrawing from retirement to pay off the debt immediately, while Claude, ChatGPT, and Perplexity all recommend against withdrawal and treating it as a last resort.
Claude
no - exhaust alternatives first
ChatGPT
no - last resort only
Gemini
yes - pay off immediately
Mistral
n/a - not provided
Perplexity
no - alternatives strongly preferred
Grok
n/a - not provided
4 successful provider outputs recorded for this question.
11. decision
Should a solo freelance developer in France register as micro-entreprise or set up a SASU?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
micro-entreprise first, switch later
ChatGPT
micro-entreprise for simplicity, SASU for growth
Gemini
micro-entreprise first, switch later
Mistral
not present
Perplexity
micro-entreprise unless high revenue
Grok
not present
4 successful provider outputs recorded for this question.
12. decision
Is it worth buying a 10-year-old used BMW with 150,000 km, or is the maintenance risk too high?
PARTIAL62/100
Recorded divergence: Claude and ChatGPT leave the door open conditionally (clean inspection, low price, repair buffer), while Gemini and Perplexity recommend against it more categorically regardless of condition.
Claude
depends - avoid unless very cheap and clean
ChatGPT
depends - possible if well-maintained
Gemini
no - high risk outweighs savings
Mistral
no stance - not present
Perplexity
no - generally not worth it
Grok
no stance - not present
4 successful provider outputs recorded for this question.
13. decision
Should I accept a job offer with 20% higher pay but a 90-minute commute each way?
PARTIAL68/100
Recorded divergence: Mistral uniquely recommends trying the job for 1-2 months before deciding, while others lean toward declining outright or negotiating alternatives without a trial period.
Claude
no – negotiate hybrid first
ChatGPT
depends – weigh costs and growth
Gemini
no – decline outright
Mistral
depends – trial it 1-2 months
Perplexity
no – decline firmly
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
14. prediction
Is nuclear power a safe and good choice for reducing carbon emissions?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
yes, important climate tool
ChatGPT
yes, part of balanced mix
Gemini
yes, safe and clean
Mistral
yes, combine with renewables
Perplexity
yes, necessary but not sole solution
Grok
N/A
5 successful provider outputs recorded for this question.
15. prediction
Will the EUR/USD exchange rate likely be higher or lower in 12 months?
SPLIT30/100
Recorded divergence: Claude, ChatGPT, and Perplexity lean toward EUR/USD being lower (USD stronger), while Gemini and Mistral explicitly recommend expecting EUR/USD to be higher driven by narrowing rate differentials.
Claude
no - USD stronger, lower EUR/USD
ChatGPT
no - lower if trends continue
Gemini
yes - marginally higher EUR/USD
Mistral
yes - slightly higher 1.10-1.15
Perplexity
no - lower ~1.00-1.09
Grok
n/a - not provided
5 successful provider outputs recorded for this question.
16. health
Should I take daily low-dose aspirin to prevent heart disease if I'm 50 with no heart problems?
AGREE91/100
No material bottom-line divergence was recorded by the judge.
Claude
no – consult doctor first
ChatGPT
no – consult doctor
Gemini
no – focus on lifestyle
Mistral
no – unless ≥10% CVD risk
Perplexity
no – don't start alone
Grok
n/a
5 successful provider outputs recorded for this question.
17. health
Is it safe to drink alcohol while taking metronidazole?
PARTIAL78/100
Recorded divergence: All agree alcohol must be avoided, but they differ on the post-treatment waiting period: Claude and ChatGPT and Mistral say 48 hours, Gemini says 72 hours, and Perplexity says 2-3 days, which is a materially different safety caveat a user would act on.
Claude
no, avoid; wait 48hrs
ChatGPT
no, avoid; wait 48hrs
Gemini
no, avoid; wait 72hrs
Mistral
no, avoid; wait 48hrs
Perplexity
no, avoid; wait 2-3 days
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
18. health
My 2-year-old has a fever of 39°C. Should I go to the emergency room now?
AGREE91/100
No material bottom-line divergence was recorded by the judge.
Claude
depends – call pediatrician first
ChatGPT
depends – monitor, call pediatrician
Gemini
depends – call pediatrician now
Mistral
depends – call pediatrician, ER if warning signs
Perplexity
depends – call pediatrician, not ER unless serious
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
19. health
Are creatine supplements safe for long-term daily use?
AGREE96/100
No material bottom-line divergence was recorded by the judge.
Claude
yes, safe 3-5g/day
ChatGPT
yes, safe recommended doses
Gemini
yes, safe long-term
Mistral
yes, safe low-risk
Perplexity
yes, safe healthy adults
Grok
N/A
5 successful provider outputs recorded for this question.
20. health
Should a 45-year-old man get a PSA test for prostate cancer screening?
PARTIAL68/100
Recorded divergence: Perplexity uniquely endorses a baseline PSA at 45 for all men as potentially useful, while Gemini leans more firmly against routine screening, and others strictly condition testing on high-risk status only.
Claude
depends – high-risk start now
ChatGPT
depends – consult doctor first
Gemini
no – unless high-risk
Mistral
depends – wait if average-risk
Perplexity
yes – baseline useful for all
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
21. health
Is a ketogenic diet safe for someone with high cholesterol?
PARTIAL55/100
Recorded divergence: Claude and ChatGPT say 'it depends, consult doctor and monitor,' while Gemini and Perplexity lean toward 'generally avoid/not recommended,' representing meaningfully different default stances on whether to proceed at all.
Claude
depends - consult doctor, monitor closely
ChatGPT
depends - consult doctor first
Gemini
no - generally not recommended
Mistral
N/A - not provided
Perplexity
no - generally not safe, avoid
Grok
N/A - not provided
4 successful provider outputs recorded for this question.
22. legal
In France, can I legally record a phone call without telling the other person?
PARTIAL78/100
Recorded divergence: Mistral introduces a notable caveat that recording for personal use may receive judicial leniency, and Perplexity highlights a court exception allowing unauthorized recordings as evidence in rare litigation cases, while Claude and Gemini give a cleaner 'always get consent' bottom line.
Claude
no - always get consent
ChatGPT
not assessed
Gemini
no - always get consent
Mistral
no - but personal use may be lenient
Perplexity
no - but rare court exceptions exist
Grok
not assessed
4 successful provider outputs recorded for this question.
23. legal
If I find cash on the street, am I legally allowed to keep it?
PARTIAL72/100
Recorded divergence: Mistral uniquely suggests small amounts (e.g., $20 or less) are generally safe to keep without reporting, while all others recommend turning in cash regardless of amount.
Claude
no - turn in first
ChatGPT
no - turn in first
Gemini
no - turn in immediately
Mistral
depends - small amounts keep, large amounts report
Perplexity
no - report to police
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
24. legal
Can my homeowners association legally fine me for parking in my own driveway?
AGREE82/100
Recorded divergence: Perplexity emphasizes HOAs generally cannot fine for private driveway parking unless CC&Rs explicitly allow it, while Gemini leans toward HOAs likely can, but all converge on the same action: check your CC&Rs.
Claude
depends, review CC&Rs
ChatGPT
yes if CC&Rs allow, check them
Gemini
yes likely, review CC&Rs
Mistral
depends on CC&Rs, check them
Perplexity
generally no unless CC&Rs say so
Grok
N/A
5 successful provider outputs recorded for this question.
25. legal
Is it legal to fly a drone over my neighbor's property in France?
PARTIAL82/100
Recorded divergence: Mistral uniquely suggests flying above 150m as 'typically safer', which contradicts standard DGAC rules and the other models' blanket advice to avoid flying over neighbor's property entirely.
Claude
no - get explicit consent first
ChatGPT
no - avoid without explicit consent
Gemini
no - stay within own land boundaries
Mistral
no - but implies 150m altitude workaround
Perplexity
no - explicit permission required
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
26. legal
Can my employer in France legally read personal emails I sent from the work computer?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
no - label personal emails clearly
ChatGPT
no - mark emails personal
Gemini
no - but label clearly
Mistral
N/A - not provided
Perplexity
no - use personal account
Grok
N/A - not provided
4 successful provider outputs recorded for this question.
27. finance
Should I buy or rent a home if I plan to stay only 4 years?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
no — rent, 5-7yr rule
ChatGPT
no — rent, flexibility wins
Gemini
no — rent, costs too high
Mistral
no — not evaluated
Perplexity
no — rent, breakeven too long
Grok
no — not evaluated
4 successful provider outputs recorded for this question.
28. finance
Is it smart to put 10% of my retirement portfolio into Bitcoin?
PARTIAL62/100
Recorded divergence: Claude conditionally endorses 5-10% for younger investors while others cap recommendations at 1-5% maximum, making Claude meaningfully more permissive than the consensus
Claude
depends - 5-10% okay for young
ChatGPT
no - start with 1-5% max
Gemini
no - limit to 1-2% only
Mistral
no - limit to 1-5%
Perplexity
no - limit to 1-3%
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
29. finance
Should I take a loan from my 401k to fund a home down payment?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
no - explore alternatives first
ChatGPT
no - last resort only
Gemini
no - exhaust other options
Mistral
depends - job stability required
Perplexity
no - alternatives first
Grok
no data provided
5 successful provider outputs recorded for this question.
30. finance
Is whole life insurance a good investment?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
no – buy term instead
ChatGPT
depends – consider alternatives
Gemini
N/A – not provided
Mistral
no – buy term instead
Perplexity
no – term plus investing better
4 successful provider outputs recorded for this question.
31. finance
Should I pay for my child's college, or have them take student loans?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - hybrid approach, retirement first
ChatGPT
depends - balanced contribution, retirement first
Gemini
depends - hybrid approach, retirement first
Mistral
depends - pay what you can, retirement first
Perplexity
depends - pay if possible, retirement first
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
32. decision
Should I quit my stable job to start my own business at 35 with two kids?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
no - build while employed
ChatGPT
conditional yes - needs plan/cushion
Gemini
no - validate on side first
Mistral
n/a - not provided
Perplexity
no - side project first
Grok
n/a - not provided
4 successful provider outputs recorded for this question.
33. decision
Is an MBA worth the cost for a mid-career software engineer?
PARTIAL72/100
Recorded divergence: Gemini gives a hard 'do not pursue' while Claude, ChatGPT, and Perplexity leave meaningful room for career-transition scenarios where an MBA is justified.
Claude
depends — only specific transitions
ChatGPT
depends — management/business goals
Gemini
no — rarely worth it
Mistral
n/a — not provided
Perplexity
no — only if 100% leaving engineering
Grok
n/a — not provided
4 successful provider outputs recorded for this question.
34. decision
Should I relocate my family to another country for a 30% salary increase?
AGREE82/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - run real numbers first
ChatGPT
depends - analyze all factors thoroughly
Gemini
depends - no without deeper analysis
Mistral
depends - verify cost of living and fit
Perplexity
depends - calculate net disposable income first
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
35. decision
Should I tell my manager I'm interviewing elsewhere to try to get a counteroffer?
PARTIAL72/100
Recorded divergence: Perplexity leans hardest against counteroffers entirely and advises negotiating before job searching, while others conditionally allow disclosing if you have a real offer and are prepared to leave.
Claude
no, wait for offer first
ChatGPT
depends, have offer first
Gemini
no, only with concrete offer
Mistral
depends, be ready to leave
Perplexity
no, avoid counteroffers entirely
Grok
N/A
5 successful provider outputs recorded for this question.
36. decision
In tech, is it better to specialize deeply or stay a generalist?
PARTIAL62/100
Recorded divergence: Gemini recommends specialization as the clear default, while Claude, Mistral, and Perplexity advocate for T-shaped skills as the ideal, and ChatGPT defers entirely to personal preference without a directional lean.
Claude
depends + T-shaped is ideal
ChatGPT
depends + purely personal choice
Gemini
yes + specialize as default
Mistral
depends + T-shaped hybrid best
Perplexity
depends + T-shaped recommended
Grok
N/A
5 successful provider outputs recorded for this question.
37. decision
Should couples merge all their finances after marriage?
AGREE91/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - hybrid system recommended
ChatGPT
depends - hybrid system recommended
Gemini
no - hybrid system recommended
Mistral
depends - hybrid system recommended
Perplexity
depends - hybrid system recommended
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
38. decision
Is it a red flag if my partner wants to check my phone?
PARTIAL62/100
Recorded divergence: Gemini gives a firm 'do not allow it' recommendation while Claude and ChatGPT emphasize context-dependence and mutual discussion as potentially legitimate exceptions.
Claude
depends – context and mutuality matter
ChatGPT
yes – but context and open talk
Gemini
yes – do not allow it
Mistral
n/a
Perplexity
yes – red flag, discuss but enforce privacy
Grok
n/a
4 successful provider outputs recorded for this question.
39. prediction
Will remote work still be common in 5 years, or will offices win out?
AGREE94/100
No material bottom-line divergence was recorded by the judge.
Claude
yes, hybrid dominates
ChatGPT
yes, hybrid becomes norm
Gemini
yes, hybrid dominates
Mistral
N/A not provided
Perplexity
yes, hybrid dominates
Grok
N/A not provided
4 successful provider outputs recorded for this question.
40. prediction
Will electric cars fully replace gas cars by 2035?
AGREE88/100
Recorded divergence: Minor variance on projected EV new-sales share by 2035 (Mistral says 60-80%, Claude says 30-50%), but all agree full replacement will not happen
Claude
no, decades beyond 2035
ChatGPT
no, challenges remain
Gemini
no, gas still prevalent
Mistral
no, but EVs dominate new sales
Perplexity
no, only new sales targeted
Grok
no data provided
5 successful provider outputs recorded for this question.
41. prediction
Will AI cause mass unemployment within the next decade?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
no - disruption not unemployment
ChatGPT
no - shifts not mass unemployment
Gemini
no - transformation not scarcity
Mistral
no - shifts not collapse
Perplexity
no - reshape not eliminate
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
42. prediction
Will housing prices in major cities keep rising over the next 5 years?
PARTIAL52/100
Recorded divergence: Gemini expects stagnation or minor declines as the base case, while Perplexity and ChatGPT lean toward continued appreciation (even if slower), creating a meaningful split on the directional outlook.
Claude
depends – uneven, market-specific
ChatGPT
yes – likely rising but slower
Gemini
no – slowdown or modest declines
Mistral
depends – slower growth, city-specific
Perplexity
yes – continued appreciation, slower pace
Grok
N/A
5 successful provider outputs recorded for this question.
43. consumer
Should I buy a used Tesla or a new Toyota if reliability matters most?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - buy new Toyota
ChatGPT
yes - buy new Toyota
Gemini
yes - buy new Toyota
Mistral
yes - buy new Toyota, minor used-Tesla caveat
Perplexity
yes - buy new Toyota
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
44. consumer
Is it worth paying for a VPN for everyday personal privacy?
PARTIAL62/100
Recorded divergence: Perplexity leans toward a clear 'yes, worth it for most people' while Gemini and Claude lean toward 'probably not necessary unless specific use cases apply'
Claude
depends - only for specific use cases
ChatGPT
depends - worth it if privacy is priority
Gemini
no - not essential for most users
Mistral
n/a - not provided
Perplexity
yes - worth it for most people
Grok
n/a - not provided
4 successful provider outputs recorded for this question.
45. consumer
Should a small business build its app natively or with a cross-platform framework?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - cross-platform default
ChatGPT
yes - cross-platform default
Gemini
yes - cross-platform default
Mistral
not provided
Perplexity
yes - cross-platform default
Grok
not provided
4 successful provider outputs recorded for this question.
46. health
Should I take vitamin D supplements year-round?
PARTIAL65/100
Recorded divergence: Perplexity leans toward winter-only supplementation as the default for most healthy adults, while Claude and Mistral lean toward year-round supplementation for most people with indoor/northern lifestyles, creating a materially different default action.
Claude
yes, year-round for most indoors/northern
ChatGPT
depends, test first consult doctor
Gemini
depends, winter priority test first
Mistral
depends, test-driven dose adjustment
Perplexity
no, winter-only unless risk factors
Grok
N/A
5 successful provider outputs recorded for this question.
47. health
Is it safe to exercise with a mild cold?
AGREE96/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - light exercise, above-neck rule
ChatGPT
yes - light exercise, above-neck rule
Gemini
yes - gentle exercise, above-neck rule
Mistral
yes - low-moderate intensity, above-neck rule
Perplexity
yes - light exercise, above-neck rule
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
48. health
Should I get my wisdom teeth removed if they aren't causing pain?
PARTIAL72/100
Recorded divergence: Gemini leans toward 'no removal needed if asymptomatic' as a default, while Claude and Mistral lean toward proactive removal if any concern exists, especially in younger patients.
5 successful provider outputs recorded for this question.
49. health
Is melatonin safe to take every night for sleep?
PARTIAL72/100
Recorded divergence: Perplexity uniquely raises a 2025 cardiovascular risk study suggesting potential heart failure risk, while Gemini is notably more discouraging ('generally not recommended') versus others framing it as cautiously acceptable short-term.
Claude
depends – short-term ok, long-term consult doctor
ChatGPT
depends – consult doctor for nightly extended use
Gemini
no – not recommended, consult doctor
Mistral
depends – short-term ok, address root causes
Perplexity
no – not proven safe, cardiac risk concern
Grok
N/A – not provided
5 successful provider outputs recorded for this question.
50. health
Should I start statins if my cholesterol is only borderline high?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
depends – lifestyle first, risk-stratify
ChatGPT
depends – lifestyle first, consult doctor
Gemini
depends – lifestyle first, consult doctor
Mistral
depends – lifestyle first, risk-stratify
Perplexity
depends – lifestyle first, risk-stratify
Grok
N/A
5 successful provider outputs recorded for this question.
51. health
Is it safe to take ibuprofen and acetaminophen at the same time?
SPLIT35/100
Recorded divergence: Gemini explicitly says it is NOT recommended to take them at the same time and advises spacing them out, while Claude, ChatGPT, Mistral, and Perplexity all say it IS generally safe to take them together simultaneously.
Claude
yes, safe to combine
ChatGPT
yes, safe to combine
Gemini
no, not recommended simultaneously
Mistral
yes, safe to combine
Perplexity
yes, safe to combine
Grok
N/A
5 successful provider outputs recorded for this question.
52. legal
Can I be fired in France for a social-media post made on my personal time?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - protected but exceptions apply
ChatGPT
yes - possible if harmful/misconduct
Gemini
yes - possible if serious misconduct
Mistral
not reviewed
Perplexity
yes - possible but privacy settings matter
Grok
not reviewed
4 successful provider outputs recorded for this question.
53. legal
Is it legal to resell concert tickets above face value?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - check local laws
ChatGPT
depends - check local laws
Gemini
generally legal with exceptions
Mistral
not assessed
Perplexity
depends - jurisdiction specific
Grok
not assessed
4 successful provider outputs recorded for this question.
54. legal
Can my landlord enter my apartment without notice in France?
AGREE88/100
Recorded divergence: Minor differences in specified notice periods (24hrs vs 48hrs vs unspecified) but all agree landlord cannot enter without notice except emergencies
Claude
no, except emergencies
ChatGPT
no, except emergencies
Gemini
no, consent required
Mistral
no, 24hr notice required
Perplexity
no, 48hr notice required
Grok
N/A
5 successful provider outputs recorded for this question.
55. legal
Am I liable if my dog bites someone who entered my yard uninvited?
PARTIAL62/100
Recorded divergence: Mistral and Perplexity lean toward 'likely not liable' for trespassers, while Claude and ChatGPT emphasize you're 'not automatically off the hook' and may still face significant liability.
Claude
depends, not off the hook
ChatGPT
depends, possible liability remains
Gemini
n/a
Mistral
likely not liable unless reckless
Perplexity
generally not liable unless negligent
Grok
n/a
4 successful provider outputs recorded for this question.
56. legal
Can I legally refuse a breathalyzer test during a traffic stop in France?
PARTIAL55/100
Recorded divergence: Perplexity distinguishes between the initial screening test (éthylotest, refusable) and the verification test (éthylomètre, not refusable), while Claude, ChatGPT, and Mistral give a blanket 'you cannot refuse' answer with no such distinction.
Claude
no - refusal illegal, criminal offense
ChatGPT
no - refusal illegal, comply
Gemini
no response provided
Mistral
no - refusal illegal, severe penalties
Perplexity
depends - initial test refusable, verification test not
Grok
no response provided
4 successful provider outputs recorded for this question.
57. finance
Should I invest in individual stocks or only index funds?
PARTIAL78/100
Recorded divergence: Models agree on index funds as primary vehicle but meaningfully differ on how much to allocate to individual stocks: Claude suggests up to 20%, Mistral up to 10%, Perplexity caps at 5%, while Gemini recommends index funds exclusively.
Claude
yes index funds, 10-20% stocks ok
ChatGPT
yes index funds, small portion stocks ok
Gemini
yes index funds exclusively
Mistral
yes index funds, under 10% stocks ok
Perplexity
yes index funds, max 5% stocks optional
Grok
N/A
5 successful provider outputs recorded for this question.
58. finance
Is it worth paying points to lower my mortgage interest rate?
AGREE97/100
No material bottom-line divergence was recorded by the judge.
Claude
depends on break-even timeline
ChatGPT
depends on break-even timeline
Gemini
depends on break-even timeline
Mistral
not included
Perplexity
depends on break-even timeline
Grok
not included
4 successful provider outputs recorded for this question.
59. finance
Should I claim Social Security at 62 or wait until 67?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends – health/finances determine timing
ChatGPT
depends – wait if healthy/stable finances
Gemini
yes – wait until 67 if possible
Mistral
N/A – not provided
Perplexity
yes – wait until 67 or 70
Grok
N/A – not provided
4 successful provider outputs recorded for this question.
60. finance
For a new business, is it better to lease equipment or buy it?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - lease for new businesses
ChatGPT
yes - lease to preserve cash
Gemini
yes - lease is better
Mistral
yes - lease for most startups
Perplexity
depends - lease now buy long-term
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
61. finance
Should I keep 6 months of expenses in savings, or invest most of it?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - 3-6 months then invest
ChatGPT
depends - 3-6 months then invest
Gemini
yes - full 6 months then invest
Mistral
depends - 3-6 months then invest
Perplexity
depends - 3-6 months then invest
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
62. decision
Should I go back to school at 40 to change careers?
PARTIAL72/100
Recorded divergence: Perplexity leans affirmatively yes while others maintain a conditional depends framing, though all agree alternatives should be considered
Claude
depends - research field requirements first
ChatGPT
depends - if benefits outweigh costs
Gemini
yes - if well-researched and purposeful
Mistral
depends - if goals and finances align
Perplexity
yes - often smart investment in growth fields
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
63. decision
Is it better to rent in the city or buy in the suburbs with a long commute?
PARTIAL62/100
Recorded divergence: Perplexity leans toward city renting as the default recommendation while others present it as a genuine depends-on-your-situation choice with no clear default winner.
Claude
depends - run the numbers
ChatGPT
depends - lifestyle vs investment
Gemini
depends - personal priorities
Mistral
depends - rent if commute over 45min
Perplexity
yes - rent city is default better
Grok
N/A
5 successful provider outputs recorded for this question.
64. decision
Should I lend money to a family member who asks?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - only if affordable/gift-mindset
ChatGPT
depends - only if affordable/gift-mindset
Gemini
depends - only if affordable/gift-mindset
Mistral
not evaluated
Perplexity
depends - only if affordable/gift-mindset
Grok
not evaluated
4 successful provider outputs recorded for this question.
65. decision
Should I have children before 30, or focus on my career first?
AGREE82/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - readiness over timing
ChatGPT
depends - personal circumstances decide
Gemini
depends - values and stability matter
Mistral
depends - not applicable/not provided
Perplexity
depends - readiness over calendar
Grok
depends - not applicable/not provided
4 successful provider outputs recorded for this question.
66. decision
Should I confront a coworker who took credit for my work, or go straight to HR?
SPLIT35/100
Recorded divergence: Gemini recommends going straight to HR while all other models recommend direct confrontation with the coworker first
Claude
yes - confront coworker first
ChatGPT
yes - confront coworker first
Gemini
no - go straight to HR
Mistral
yes - confront coworker first
Perplexity
yes - confront coworker first
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
67. decision
Should I sell my house now or wait, hoping prices rise?
PARTIAL72/100
Recorded divergence: Mistral and Perplexity lean more toward selling now citing near-peak prices and stalling appreciation, while Claude and ChatGPT explicitly avoid a directional recommendation and emphasize personal circumstances over market timing.
Claude
depends - personal circumstances first
ChatGPT
depends - assess situation and trends
Gemini
N/A - not provided
Mistral
lean sell now - prices near peak
Perplexity
lean sell now - minimal upside waiting
Grok
N/A - not provided
4 successful provider outputs recorded for this question.
68. decision
Is it worth hiring a personal trainer, or can I get results alone?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - situation-specific guidance
ChatGPT
depends - motivation and knowledge key
Gemini
yes - at least short-term
Mistral
depends - try a few sessions first
Perplexity
yes - especially for faster results
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
69. prediction
Will interest rates be higher or lower in 2 years?
AGREE82/100
No material bottom-line divergence was recorded by the judge.
Claude
yes - moderately lower rates
ChatGPT
depends - stay informed, diversify
Gemini
yes - gradually lower rates
Mistral
yes - lower, stay flexible
Perplexity
yes - slightly lower rates
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
70. prediction
Will a college degree still be worth the cost in 10 years?
AGREE88/100
No material bottom-line divergence was recorded by the judge.
Claude
depends on field/cost
ChatGPT
depends on field/goals
Gemini
yes with strategic choices
Mistral
not reviewed
Perplexity
yes with major/debt caveats
Grok
not reviewed
4 successful provider outputs recorded for this question.
71. prediction
Will social-media use among teenagers decline in the next 5 years?
AGREE92/100
No material bottom-line divergence was recorded by the judge.
Claude
no - shifts not declines
ChatGPT
no - platform shifts expected
Gemini
no - use stays high or increases
Mistral
n/a - not provided
Perplexity
no - heavy use continues
Grok
n/a - not provided
4 successful provider outputs recorded for this question.
72. prediction
Will renewable energy be cheaper than fossil fuels everywhere by 2030?
AGREE91/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - not universally, most places
ChatGPT
depends - many but not everywhere
Gemini
N/A - not provided
Mistral
depends - most regions not all
Perplexity
depends - dominant but not everywhere
Grok
N/A - not provided
4 successful provider outputs recorded for this question.
73. consumer
Should I upgrade to the latest iPhone every year?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
no - upgrade every 2-3 years
ChatGPT
no - upgrade every 2-3 years
Gemini
no - upgrade every 2-3 years
Mistral
no - upgrade every 2-3 years
Perplexity
no - upgrade every 2-3 years
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
74. consumer
Is it worth buying an extended warranty on a new car?
AGREE82/100
No material bottom-line divergence was recorded by the judge.
Claude
no - self-insure instead
ChatGPT
depends - keep car long-term
Gemini
no - save money separately
Mistral
no - with noted exceptions
Perplexity
no - unless 10+ year owner
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
75. consumer
Should I use a robo-advisor or a human financial advisor?
AGREE95/100
No material bottom-line divergence was recorded by the judge.
Claude
depends - complexity/wealth threshold
ChatGPT
depends - complexity threshold
Gemini
depends - complexity threshold
Mistral
depends - complexity threshold
Perplexity
depends - complexity/hybrid approach
Grok
N/A - not provided
5 successful provider outputs recorded for this question.
Examples: a panel on a question that has a checkable answer
Representative scenarios, not logged transcripts. Each shows how a multi-model panel handles a question with a verifiable answer, which we then check against a primary source. Separate from the 75-question set above. The patterns: a panel catching one model's factual error instead of averaging it in; a short factual question the panel simply gets right; and a tool-using agent declining to fabricate an answer it cannot actually know.
decision
AWS, Google Cloud or Azure to launch a B2B healthcare SaaS in Europe (GDPR + French health-data hosting, tight budget, due-diligence ready in a year)?
What the panel recorded
Example. A panel compares the three hyperscalers across five criteria and leans to Azure, with AWS as the technical alternative. The models do not fully agree: one asserts that AWS and Google Cloud do not hold French HDS (Hébergement de Données de Santé) certification. A good synthesis flags that claim as contested rather than passing it through.
All three hyperscalers hold French HDS certification (AWS eu-west-3, Azure France Central/South, Google Cloud europe-west9). So the choice is not gated by 'who is certified' — it is a fit question: contracting terms and BAA/DPA maturity, whether the health authority or a hospital buyer expects a specific provider, data-residency wording, and cost. Azure is a defensible default for a French health SaaS on standardised health agreements, but AWS and Google Cloud are equally eligible on the certification point one model got wrong.
Which models were right
The panel's overall Azure lean was reasonable, and the synthesis was right not to trust the 'AWS/GCP are not HDS-certified' claim — the ANS registry contradicts it directly. It is a case of the panel catching one model's factual error instead of averaging it in.
Limits of this case
Certification status and regional coverage change; the ANS registry is the live source. Certification is a floor, not the whole compliance question — contracts, configuration and the specific buyer's requirements still decide it.
consumer
Can I buy tickets for the Sagrada Família towers on site, since they were sold out online?
What the panel recorded
Example. Asked this, a panel answers directly: no — there is no physical ticket office, everything is sold through sagradafamilia.org, and in 2026 (Gaudí centenary) tower tickets typically go 2–3 weeks ahead. It adds practical steps: watch the site for last-minute cancellations, try several dates, avoid third-party resellers, fall back to the basic entry without towers.
Correct. No on-site sales; tower access is a separate ticket from general admission; it routinely sells out weeks ahead from May to October. The only realistic on-the-day route is refreshing the official site for released cancellations. Third-party 'tower ticket' sites are resellers, not the basilica.
Which models were right
Cove's answer matched the official policy exactly, including the detail that towers are a separate ticket. A short factual question where the panel simply got it right, and the primary source confirms it.
Limits of this case
Operating policy can change — the official site is the source of truth on the day. Prices and lead times move with the season and the centenary programme.
consumer
Find Amazon listings with an active coupon for a 40–60 W single-port USB-C charger and a 10 ft 60 W cable.
What the panel recorded
Example, tool-using agent. It can return verifiable cable options (specific listings with prices) but should refuse to produce a table of 'active coupons': Amazon coupons are personalised and change in real time, so a scrape cannot establish that a coupon is live for a given shopper. The honest output is the verifiable part plus the check-it-yourself route (the 'Coupons' filter in the results sidebar).
Primary sources
Amazon — About Amazon coupons“Coupons are offered at Amazon's discretion, can be personalised to the shopper, and are subject to change or removal at any time; clip a coupon to apply it at checkout.”
Verified answer
The refusal was the right call. Coupon availability on Amazon is account-specific and time-varying, and is not reliably readable from a public page scrape — a claimed 'active coupon' table would have looked authoritative and been wrong for many readers. Returning the verifiable part (specific listings and prices) plus the exact step to confirm a coupon on your own account is the honest answer.
Which models were right
This case is not about a factual split — it is about a model declining to fabricate. It belongs here because the disagreement study's whole point is that a confident, well-formatted answer is not the same as a true one; sometimes the correct move is to say what cannot be known.
Limits of this case
Specific prices and listings date quickly. The durable point is the boundary — an agent saying what it cannot verify — not any particular product.
What the study can—and cannot—show
It shows that materially different recommendations occurred often in this selected sample, and that a single polished answer can hide those differences. It does not show which answer was correct, that consensus improves accuracy, or that the reported percentages generalize beyond these 75 questions.
One failure mode this study can't rule out: models sharing a blind spot. If every model in the panel was trained on the same stale or incomplete information, they can converge on the same wrong answer — agreement, in that case, is not evidence of correctness, just evidence that the training data was uneven in the same direction for everyone asked. The 45 questions classified AGREE above still warrant a primary-source check on any claim consequential enough to act on.
Methodology exploratory-v1 · Published 2026-06-25 · JSON SHA-256 b3be0e91aa3e336b5c17411859e89b55e337ffeedfc1865986d9f0fea741a28c
This study was run by the team behind Satcove, a product of Abyssal Group, a data-intelligence company. Satcove asks your question to six AI models at once and returns one verdict with the agreements, the conflicts and an agreement score. Why multi-model consensus, and where the idea comes from, is in our story.