/

Guide

How we score AI companion apps

Five criteria with fixed weights, two calibration anchors, and a stated research basis on every profile. This is the scorecard behind every number on the site.

By AI Companion Radar Editors

· 4 min read

Every app on AI Companion Radar is scored out of ten on the same five criteria with the same weights. The number is recalculated whenever we re-test, and every profile shows the date of its last review. This page is the scorecard behind that number, so you can see exactly what a 7.4 means and where you might weigh things differently.

The five criteria

Criterion

Weight

What we look at

Conversation quality

30%

Whether replies stay in character, follow instructions, avoid repetition and hold a scene over a long session. Model choice and reply length limits count here.

Memory

20%

Recall of facts, events and preferences across sessions; whether you can see and edit what the app remembers; whether memory is gated behind a higher tier or extra credits.

Media generation

20%

Image and video quality, character consistency across generations, reliability of completion, and the real cost per image or clip. Apps with no media generation are scored on the other four criteria and the result is scaled to ten.

Pricing honesty

15%

Is the price visible before sign-up? Does the renewal price match the intro price? Is there a lasting free tier? How much of the product sits behind tokens, coins or credits on top of the subscription?

Content policy clarity

15%

Whether the published terms say what adult content is permitted and what is banned, whether the marketing matches the terms, and whether age assurance is explained. A clear “not allowed” policy scores well here; “unfiltered” marketing over terms that ban explicit content scores badly.

Calibration anchors: Candy AI 8.2, OurDream AI 8.0, MyBabes 8.6

How a score is built

Each criterion gets an integer from 1 to 10. The overall score is the weighted average, rounded to one decimal. Two anchors keep the scale honest across reviewers: Candy AI sits at 8.2 (conversation 8, memory 7, media 9, pricing 6, policy 9) and OurDream AI at 8.0 (8, 8, 7, 7, 9). An app that beats both on every criterion should score above them; an app that only beats them on marketing should not.

Desk research versus hands-on testing

Every profile states its research basis. Desk-researched profiles are built from the app’s own site, pricing pages, terms, store listings and a reading of public user reports. Hands-on tested profiles were run by us on a paid account for at least a week. Where an app hides its pricing behind a login, we say so and quote what reviewers report, with the date.

What does not affect the score

  • Whether the app pays anyone. It does not pay us; we run no affiliate links and sell no placements.

  • User counts, awards and “number one” badges. We do not repeat them.

  • Traffic or popularity. A widely used app with poor memory still gets a poor memory score.

Rankings by use case

The overall ranking orders apps by their score. Use-case rankings such as best for video or best memory use the same criterion scores but only include apps where the feature exists on a published plan, and readers should treat them as a shortlist rather than a verdict.

Corrections and re-tests

Profiles are re-checked at least once a quarter, and sooner when an app changes its pricing or policy. If you spot an error, email hello@aicompanionradar.com with the app name and a source. Corrections with a source are usually live within a week and do not change scores on their own; a re-test does.