How we test AI girlfriend apps
Last reviewed: June 12, 2026 · Maintained by the GirlfriendsAI test desk
Every app on this site is paid for at our own expense, installed across the platforms it officially supports, and used for at least seven days - usually closer to two weeks. We score each one on a fifty-point rubric covering conversation, image generation, voice, and pricing honesty. The rubric is published below in full. Anything you read on this site that claims a number - "46/50 coherence," "1.2-second call latency," "18/20 fact recall" - comes from this process.
The 50-point rubric
Five categories, weighted by what we believe actually matters to users in 2026. Image and voice are weighted lower than chat because every app in the category can produce some image and some voice - the difference between a good and a great AI companion almost always lives in the conversation.
Conversation quality
Coherence across 200 messages, in-character consistency, emotional pickup, callbacks to facts shared in earlier sessions. Largest weight by design.
Image generation
Character lock across 50 generations, prompt obedience, latency, NSFW policy clarity, refusal patterns.
Voice & calls
TTS quality, call latency on a controlled Wi-Fi connection, interruption handling, voice-clone honesty.
Pricing honesty
Does the free tier work? Are paywalls disclosed before you tap them? Are coin economies explained? Refunds?
Privacy & safety
Privacy policy reading, opt-in vs opt-out on training data, data deletion path, encryption claims verified.
What we run
The 200-message coherence script
Every app gets the same controlled chat sequence, run by the same reviewer, against the same persona we build inside the app. The script is broken into five blocks of forty messages: small talk; deliberate fact-sharing (job, pet's name, hobby, fake vacation plan); a long argument we pick to stress-test in-character behaviour; an emotional curveball at message 80 ("had a rough day"); a memory test at messages 150–200 where we reference each shared fact obliquely to see whether the model recalls it.
Scoring is binary per fact: she either recalled it or she didn't. Tone is scored by the reviewer on a 0–5 scale per block. The blocks are summed, capped at 20.
The 50-image character-lock test
We request 50 image generations of the same persona across varied prompts (full body, close-up, different outfits, different backgrounds, day and night, indoors and outdoors). Two scorers independently flag any image where the model has visibly drifted from the locked face, hair, or build. The score is 10 minus drifted images, capped between 0 and 10. NSFW images are tested only on apps that advertise NSFW; refusals are noted but not penalized when they match the published policy.
The 10-call voice test
Ten voice calls, five on Wi-Fi and five on LTE, made from a single phone in a quiet room. We measure response latency (time from end-of-our-speech to start-of-her-speech) with a stopwatch; quality is scored on a 0–5 scale per call by the same reviewer.
Privacy posture reading
We read each app's privacy policy and terms of service end-to-end. We score based on (a) whether your chats train future models by default, (b) whether you can delete history, (c) whether the company has had a public incident in the last 24 months, and (d) what jurisdiction it operates from.
Test hardware & environment
All mobile testing runs on a 2024 iPhone 15 and a 2024 Pixel 8a. Desktop testing runs on a 2023 MacBook Pro M3 (16 GB RAM) over a 300 Mbps fibre connection in a single home office. Voice tests use the native phone microphone - no studio mic, because that isn't what most users have.
What we don't do
- We don't accept payment to feature an app. No "sponsored review," no "branded content," no "guest editorial." Apps without an affiliate program (currently Anima, Linky, Mate, Girlfriend GPT, AI2U, Yandere Simulator) are reviewed and ranked using the exact same rubric as apps that pay a commission, and you'll see them appearing at every rank from top to bottom of our lists.
- We don't accept early review copies in exchange for embargoes. Every review goes live after the test window closes, on our own timeline.
- We don't review apps we haven't paid for and used. If an app appears on this site, the test desk has its receipts.
- We don't promise positive coverage in exchange for affiliate-program access. Several apps we currently link to received scores below 4.0; we kept the link active because the rubric came out where it came out.
How we handle disagreement
Two reviewers score each app independently on the rubric. Discrepancies above 4 points trigger a third reviewer who re-tests in a separate session. We publish the final score, not the average - the third reviewer decides which of the first two scores stands. This catches the case where one reviewer happened to build a persona the app handles well and the other built one it handles badly.
When we re-test
Every app is re-tested in full once per quarter. Mid-quarter, we re-open a review when an app meaningfully changes (model update, pricing change, NSFW policy flip). The "Last tested" date stamped on every review tells you exactly when we last sat down with the app - if it's more than 90 days old, treat the score as directional, not current.
Where to push back
If you've used one of these apps and your experience disagrees with our score, write us. We re-test apps that get sustained pushback - Joi got moved up two places in our June 2026 round because four readers in a row pushed back on our coin-economy criticism, and on retest we agreed with them on two of three points. The methodology is meant to be transparent because it's meant to be argued with.
Try our #1 pick: OurDream Free → See the rankings Meet the test desk