Resources / Landeed Labs / Terra Black / Benchmarks

BENCHMARK

Introducing Landeed Property Intelligence Benchmark

August 2026 · 5 min read

Landeed Property Intelligence Bench

Terra Black T2
Agent Benchmark

We benchmarked Terra Black against four frontier models.
It ranked first.

100 real Indian title-verification matters. Five AI models. Terra Black scored 0.891 overall and ranked #1, ahead of every frontier model tested.

macro score
0.891
of 5 systems
#1
real matters
100

LANDEED PROPERTY INTELLIGENCE BENCH · 100 REAL INDIAN PROPERTY MATTERS

Title verification isn't just a reading test.
It takes judgment.

A title-verification report must do more than extract facts; it must connect the evidence, interpret its implications and reach a defensible conclusion. So we benchmarked 100 real Indian matters against one fixed legal rubric, applied the same way to every model.

  • 100

    real Indian title matters scored

  • 500+

    real source documents

  • 8

    legal dimensions graded

  • ~3–4 min

    per full title report

THE LEADERBOARD

Five systems. One benchmark.
Terra Black ranked first.

Across 100 title matters, Terra Black scored 0.891, the highest macro score among the five models tested.

#1 OF 5 SYSTEMS

0.891

macro score: mean criterion pass-rate across 100 title matters

MACRO SCORE: MEAN CRITERION PASS-RATE ACROSS 100 TITLE MATTERS
Terra Black0.891
Sonnet-50.863
Gemini-3.60.852
Opus-4.80.844
GPT-5.60.769

Property cases passed in full: every scored criterion satisfied

100 cases · strict all-pass

  • 26Terra Blackof 100
  • 14Sonnet-5of 100
  • 9Opus-4.8of 100
  • 8Gemini-3.6of 100
  • 1GPT-5.6of 100

The average scores are close. The all-pass test is less forgiving.

WHERE TERRA LEADS

Frontier models read well.
Terra Black goes further.

The field is strong at extracting facts from property documents. Terra Black's advantage appears where the work requires applying Indian property law and reaching a judgment, while staying competitive across the rest of the title-verification workflow.

Pass rate across eight legal dimensions · 100 real Indian title cases

Pass rate across eight legal dimensions · 100 real Indian title property cases.

DIMENSIONTerra BlackSonnet-5Gemini-3.6Opus-4.8GPT-5.6
Statutory & tenure0.8880.7570.8460.7710.682
Mortgage & security0.7970.7160.6910.7030.686
Chain of title0.8880.8800.8880.7980.810
Documents scrutinized0.9380.9380.9260.9070.895
Title opinion & recs0.8800.7700.9500.7600.320
Encumbrances & liens0.8840.9190.8490.8600.907
Parties0.9240.9310.7680.9200.720
Property identification0.8930.9170.9380.9300.909

Highest score in each dimension shown in bold. Report structure scored 1.000 for all five and is omitted as a non-differentiator.

Highest score in each dimension marked with a dot. Report structure scored 1.000 for all five and is omitted as a non-differentiator.

#1 OF ALL 5 SYSTEMS

0.888

Statutory & tenure standing

Freehold, RERA, tenure and conversion, joint-family interests and minors' shares, and what they mean for the title.

AHEAD OF THE OVERALL LEADER

0.797

Mortgage & security opinion

Assessing whether the property can support a valid and enforceable mortgage.

CASE BY CASE

The lead isn't just in the average.
It holds case by case.

Across 100 title matters, Terra Black matched or beat each model on a majority of cases and scored higher on average in every head-to-head comparison.

CASE-LEVEL OUTCOMES · 100 TITLE REVIEWS

Terra higherEqual scoreOpponent highermean delta per case →
Terra vs. GPT-5.687%13%+0.122
Terra vs. Opus-4.857%19%24%+0.048
Terra vs. Gemini-3.673%27%+0.039
Terra vs. Sonnet-551%21%28%+0.029

SPEED & QUALITY

The highest score. In about four minutes.

Terra Black produces a complete title-verification report in about four minutes in production, while posting the benchmark's highest score at 0.891. High-quality property intelligence doesn't have to mean a long wait.

BENCHMARK QUALITY × OBSERVED COMPLETION TIME

top-right is the sweet spot

  • Terra Black 3.5 min · 0.891
  • Sonnet-5 6.8 min · 0.863
  • Gemini-3.6 1.2 min · 0.852
  • Opus-4.8 ~8 min · 0.844
  • GPT-5.6 2.4 min · 0.769

Timings are approximate, derived from completion timestamps in benchmark runs.

ROBUSTNESS & SCALE

Real-world property matters, at real-world scale.

The benchmark used real title-verification matters with large, complex document sets. On one of its hardest matters, Terra Black scored 0.833, the highest of the systems tested.

Case size across the benchmark

megabytes

  • 9 MB median
  • 110 MB largest
  • 129 MB hardest

File size describes the workload, not legal difficulty.

The single hardest case (~129 MB)

score, higher is better

Terra Black0.833
Opus-4.80.722
Gemini-3.60.667
GPT-5.60.556

On the largest case by file size (~110 MB), a frontier model led and Terra Black trailed at 0.733.

METHODOLOGY & DISCLOSURE

One benchmark. Five harnesses. Same standard.

All five models were evaluated on the same 100 title matters, against the same fixed legal rubric, using the same criterion-level scoring process.

  • ONE FIXED RUBRIC

    Every title matter was scored criterion by criterion against the same legal rubric.

  • CRITERION-LEVEL SCORING

    Each requirement was scored independently, then aggregated into matter-, dimension- and benchmark-level scores.

  • SEPARATE EVALUATOR MODEL

    A separate model evaluated every output against the same rubric at temperature 0.

  • ONE OUTPUT PER TITLE MATTER

    One evaluated output per system, per title matter. No best-of-N selection.

Internal benchmark.

See what Terra does with your Property Case.

We benchmarked Terra Black on 100 real title-verification matters. Now try it on one your team knows well and compare the evidence it finds, the analysis it produces, and the opinion it reaches.

Test Terra on a property case →(opens in a new tab)