跳至内容
AiMing

Methodology

How we measure accuracy

You cannot check a prediction about the future, so we tested against the past: take charts of people whose lives are documented, hide their names, and see how well the AI reads events that already happened.

结果

86.6%

433/500 答对

我们取来 20 多个国家、100 位生平记载详尽的公众人物的命盘,全部隐去姓名、登记为 TEST-001 至 TEST-100,使模型无法凭训练资料中的传记作答,再就真实发生过的事件提出 500 个问题,逐一与史实比对。

在八字领域与通用 AI 的比较

  • AiMing92.14%
  • ChatGPT GPT-528.3%
  • Claude 4.520.61%
  • Gemini 2.5 Flash15.15%

方法: 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge (2025-10)

并非每个领域都一样准

我们连表现最差的类别也一并公布。号称各领域一样准的系统,通常是没有量测过。

  • ความเข้ากัน96.7%
  • การงาน92.1%
  • ทั่วไป88.9%
  • สุขภาพ88.9%
  • จังหวะเวลา83.1%
  • ครอบครัว82.5%
  • การเงิน82.5%
  • ความรัก77.8%

八字告诉不了你的事

  • 只能给出时间区间,无法指出确切日期。
  • 无法说出人名、公司名或具体地点。
  • 能提示健康倾向,但不作诊断——我们不是医生。
  • 八字描述的是条件与机率,不是既定的命。
  • 你的决定与行动会改变结果。

The test, step by step

  1. 01

    Pick people whose lives are documented

    We selected 100 public figures from over 20 countries whose life histories are thoroughly recorded and independently checkable — heads of state, artists, athletes, scientists.

  2. 02

    Hide every name

    Each chart was registered as TEST-001 through TEST-100 with no name attached, so the model could not answer from a biography it had read during training. This is the step most evaluations skip.

  3. 03

    Ask about things that already happened

    We asked 500 questions about real events in those lives, across 8 categories. A prediction about the future cannot be checked. The past can.

  4. 04

    Compare against the record

    An independent judge compared each answer to the documented facts, without penalising the model for saying the same thing in a different language — "ไม้", "Wood" and "木" all count.

We audited for data leakage too

Anonymisation is not airtight — occasionally the model infers who a chart belongs to. So we audited every answer and found 8.6% where it named the real person. Those cases did score higher, but the effect on the overall result was only 0.9 points. We publish this because if we didn't, you should be suspicious.

Why this number differs from others you may have seen

We have two results measuring two different things. The first, 92.14%, measures BaZi theory knowledge across 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge. The second, 86.6%, measures reading a real life — a much harder task. We show both, and label what each one measures.

The median score across the 500 questions was 95/100 while the mean was 84.7/100. That gap is informative: when the system is right it is nearly fully right, and when it is wrong it is clearly wrong. There is very little middle ground.

And all of it rests on one person: Ravi Aunyakan, who has practised for 18 years, decided what counted as a correct answer. When the system got one wrong, we sent it back to him to explain which rule it broke, and that rule went into the knowledge base.