본문으로 건너뛰기
AiMing

Methodology

How we measure accuracy

You cannot check a prediction about the future, so we tested against the past: take charts of people whose lives are documented, hide their names, and see how well the AI reads events that already happened.

결과

86.6%

433/500 정답

20개국 이상에서 생애가 상세히 기록된 유명인 100명의 사주를 모아, 이름을 모두 가리고 TEST-001부터 TEST-100까지로 등록해 학습 데이터의 전기로 답할 수 없게 했습니다. 그런 뒤 실제로 일어난 사건에 대해 500개 문항을 묻고, 기록과 하나씩 대조했습니다.

사주 영역에서 일반 AI와 비교

  • AiMing92.14%
  • ChatGPT GPT-528.3%
  • Claude 4.520.61%
  • Gemini 2.5 Flash15.15%

검증 방법: 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge (2025-10)

모든 분야가 똑같이 정확하지는 않습니다

가장 성적이 나쁜 분야도 함께 공개합니다. 모든 분야가 똑같이 정확하다고 주장하는 시스템은 대개 측정하지 않은 것입니다.

  • ความเข้ากัน96.7%
  • การงาน92.1%
  • ทั่วไป88.9%
  • สุขภาพ88.9%
  • จังหวะเวลา83.1%
  • ครอบครัว82.5%
  • การเงิน82.5%
  • ความรัก77.8%

사주가 알려 줄 수 없는 것

  • 시기의 범위는 말할 수 있어도, 확정된 날짜는 알려 줄 수 없습니다.
  • 사람 이름, 회사명, 구체적인 장소는 알려 줄 수 없습니다.
  • 건강상 경향은 짚어 주지만 진단은 하지 않습니다 — 저희는 의사가 아닙니다.
  • 사주는 조건과 확률을 말할 뿐, 정해진 운명을 말하지 않습니다.
  • 당신의 판단과 행동이 결과를 바꿉니다.

The test, step by step

  1. 01

    Pick people whose lives are documented

    We selected 100 public figures from over 20 countries whose life histories are thoroughly recorded and independently checkable — heads of state, artists, athletes, scientists.

  2. 02

    Hide every name

    Each chart was registered as TEST-001 through TEST-100 with no name attached, so the model could not answer from a biography it had read during training. This is the step most evaluations skip.

  3. 03

    Ask about things that already happened

    We asked 500 questions about real events in those lives, across 8 categories. A prediction about the future cannot be checked. The past can.

  4. 04

    Compare against the record

    An independent judge compared each answer to the documented facts, without penalising the model for saying the same thing in a different language — "ไม้", "Wood" and "木" all count.

We audited for data leakage too

Anonymisation is not airtight — occasionally the model infers who a chart belongs to. So we audited every answer and found 8.6% where it named the real person. Those cases did score higher, but the effect on the overall result was only 0.9 points. We publish this because if we didn't, you should be suspicious.

Why this number differs from others you may have seen

We have two results measuring two different things. The first, 92.14%, measures BaZi theory knowledge across 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge. The second, 86.6%, measures reading a real life — a much harder task. We show both, and label what each one measures.

The median score across the 500 questions was 95/100 while the mean was 84.7/100. That gap is informative: when the system is right it is nearly fully right, and when it is wrong it is clearly wrong. There is very little middle ground.

And all of it rests on one person: Ravi Aunyakan, who has practised for 18 years, decided what counted as a correct answer. When the system got one wrong, we sent it back to him to explain which rule it broke, and that rule went into the knowledge base.