Skip to content
AiMing

Methodology

How we measure accuracy

You cannot check a prediction about the future, so we tested against the past: take charts of people whose lives are documented, hide their names, and see how well the AI reads events that already happened.

The result

86.6%

433/500 correct

We took the charts of 100 well-known people from over 20 countries whose lives are thoroughly documented, registered them with every name hidden as TEST-001 through TEST-100 so the model could not recall them from training data, asked 500 questions about events that really happened, and compared each answer to the record.

Against general AI, on BaZi

  • AiMing92.14%
  • ChatGPT GPT-528.3%
  • Claude 4.520.61%
  • Gemini 2.5 Flash15.15%

Method: 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge (2025-10)

It is not equally accurate at everything

We publish our worst category too. A system that claims equal accuracy everywhere is a system that hasn’t been measured.

  • Compatibility96.7%
  • Career92.1%
  • General88.9%
  • Health88.9%
  • Timing83.1%
  • Family82.5%
  • Finance82.5%
  • Love77.8%

What BaZi cannot tell you

  • It gives time windows, not exact dates.
  • It cannot name a person, a company, or a specific place.
  • It flags health tendencies. It does not diagnose — we are not doctors.
  • BaZi describes conditions and probabilities, not fixed fate.
  • Your decisions and actions change the outcome.

The test, step by step

  1. 01

    Pick people whose lives are documented

    We selected 100 public figures from over 20 countries whose life histories are thoroughly recorded and independently checkable — heads of state, artists, athletes, scientists.

  2. 02

    Hide every name

    Each chart was registered as TEST-001 through TEST-100 with no name attached, so the model could not answer from a biography it had read during training. This is the step most evaluations skip.

  3. 03

    Ask about things that already happened

    We asked 500 questions about real events in those lives, across 8 categories. A prediction about the future cannot be checked. The past can.

  4. 04

    Compare against the record

    An independent judge compared each answer to the documented facts, without penalising the model for saying the same thing in a different language — "ไม้", "Wood" and "木" all count.

We audited for data leakage too

Anonymisation is not airtight — occasionally the model infers who a chart belongs to. So we audited every answer and found 8.6% where it named the real person. Those cases did score higher, but the effect on the overall result was only 0.9 points. We publish this because if we didn't, you should be suspicious.

Why this number differs from others you may have seen

We have two results measuring two different things. The first, 92.14%, measures BaZi theory knowledge across 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge. The second, 86.6%, measures reading a real life — a much harder task. We show both, and label what each one measures.

The median score across the 500 questions was 95/100 while the mean was 84.7/100. That gap is informative: when the system is right it is nearly fully right, and when it is wrong it is clearly wrong. There is very little middle ground.

And all of it rests on one person: Ravi Aunyakan, who has practised for 18 years, decided what counted as a correct answer. When the system got one wrong, we sent it back to him to explain which rule it broke, and that rule went into the knowledge base.