Methodology
How we measure accuracy
You cannot check a prediction about the future, so we tested against the past: take charts of people whose lives are documented, hide their names, and see how well the AI reads events that already happened.
The result
86.6%
433/500 correct
We took the charts of 100 well-known people from over 20 countries whose lives are thoroughly documented, registered them with every name hidden as TEST-001 through TEST-100 so the model could not recall them from training data, asked 500 questions about events that really happened, and compared each answer to the record.
Against general AI, on BaZi
- AiMing92.14%
- ChatGPT GPT-528.3%
- Claude 4.520.61%
- Gemini 2.5 Flash15.15%
Method: 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge (2025-10)
It is not equally accurate at everything
We publish our worst category too. A system that claims equal accuracy everywhere is a system that hasn’t been measured.
- Compatibility96.7%
- Career92.1%
- General88.9%
- Health88.9%
- Timing83.1%
- Family82.5%
- Finance82.5%
- Love77.8%
What BaZi cannot tell you
- It gives time windows, not exact dates.
- It cannot name a person, a company, or a specific place.
- It flags health tendencies. It does not diagnose — we are not doctors.
- BaZi describes conditions and probabilities, not fixed fate.
- Your decisions and actions change the outcome.
The test, step by step
- 01
Pick people whose lives are documented
We selected 100 public figures from over 20 countries whose life histories are thoroughly recorded and independently checkable — heads of state, artists, athletes, scientists.
- 02
Hide every name
Each chart was registered as TEST-001 through TEST-100 with no name attached, so the model could not answer from a biography it had read during training. This is the step most evaluations skip.
- 03
Ask about things that already happened
We asked 500 questions about real events in those lives, across 8 categories. A prediction about the future cannot be checked. The past can.
- 04
Compare against the record
An independent judge compared each answer to the documented facts, without penalising the model for saying the same thing in a different language — "ไม้", "Wood" and "木" all count.
We audited for data leakage too
Anonymisation is not airtight — occasionally the model infers who a chart belongs to. So we audited every answer and found 8.6% where it named the real person. Those cases did score higher, but the effect on the overall result was only 0.9 points. We publish this because if we didn't, you should be suspicious.
Why this number differs from others you may have seen
We have two results measuring two different things. The first, 92.14%, measures BaZi theory knowledge across 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge. The second, 86.6%, measures reading a real life — a much harder task. We show both, and label what each one measures.
The median score across the 500 questions was 95/100 while the mean was 84.7/100. That gap is informative: when the system is right it is nearly fully right, and when it is wrong it is clearly wrong. There is very little middle ground.
And all of it rests on one person: Ravi Aunyakan, who has practised for 18 years, decided what counted as a correct answer. When the system got one wrong, we sent it back to him to explain which rule it broke, and that rule went into the knowledge base.