本文へスキップ
AiMing

Methodology

How we measure accuracy

You cannot check a prediction about the future, so we tested against the past: take charts of people whose lives are documented, hide their names, and see how well the AI reads events that already happened.

結果

86.6%

433/500 正答

20か国以上、生涯の記録が詳細に残る著名人100名の命式を用意し、氏名をすべて伏せて TEST-001〜TEST-100 として登録することで、学習データ中の伝記から答えられないようにしました。そのうえで実際に起きた出来事について500問を出題し、記録と一つずつ照合しています。

四柱推命における一般的なAIとの比較

  • AiMing92.14%
  • ChatGPT GPT-528.3%
  • Claude 4.520.61%
  • Gemini 2.5 Flash15.15%

検証方法: 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge (2025-10)

すべての分野で同じ精度ではありません

もっとも成績の悪い分野も公開します。どの分野も同じ精度だと謳うシステムは、たいてい測っていないだけです。

  • ความเข้ากัน96.7%
  • การงาน92.1%
  • ทั่วไป88.9%
  • สุขภาพ88.9%
  • จังหวะเวลา83.1%
  • ครอบครัว82.5%
  • การเงิน82.5%
  • ความรัก77.8%

四柱推命にわからないこと

  • 時期の幅は示せますが、確定した日付は出せません。
  • 人名・社名・具体的な地名は出せません。
  • 健康上の傾向は示しますが、診断はしません——医師ではありません。
  • 四柱推命が示すのは条件と確率であって、定まった運命ではありません。
  • あなたの判断と行動が結果を変えます。

The test, step by step

  1. 01

    Pick people whose lives are documented

    We selected 100 public figures from over 20 countries whose life histories are thoroughly recorded and independently checkable — heads of state, artists, athletes, scientists.

  2. 02

    Hide every name

    Each chart was registered as TEST-001 through TEST-100 with no name attached, so the model could not answer from a biography it had read during training. This is the step most evaluations skip.

  3. 03

    Ask about things that already happened

    We asked 500 questions about real events in those lives, across 8 categories. A prediction about the future cannot be checked. The past can.

  4. 04

    Compare against the record

    An independent judge compared each answer to the documented facts, without penalising the model for saying the same thing in a different language — "ไม้", "Wood" and "木" all count.

We audited for data leakage too

Anonymisation is not airtight — occasionally the model infers who a chart belongs to. So we audited every answer and found 8.6% where it named the real person. Those cases did score higher, but the effect on the overall result was only 0.9 points. We publish this because if we didn't, you should be suspicious.

Why this number differs from others you may have seen

We have two results measuring two different things. The first, 92.14%, measures BaZi theory knowledge across 10,000 questions across 100 charts, master-verified ground truth, LLM-as-judge. The second, 86.6%, measures reading a real life — a much harder task. We show both, and label what each one measures.

The median score across the 500 questions was 95/100 while the mean was 84.7/100. That gap is informative: when the system is right it is nearly fully right, and when it is wrong it is clearly wrong. There is very little middle ground.

And all of it rests on one person: Ravi Aunyakan, who has practised for 18 years, decided what counted as a correct answer. When the system got one wrong, we sent it back to him to explain which rule it broke, and that rule went into the knowledge base.