Resources / Landeed Labs / Terra Black / Benchmarks
BENCHMARK
Introducing Landeed Property Intelligence Benchmark
August 2026 · 5 min read
Landeed Property Intelligence Bench
Terra Black T2
Agent Benchmark
We benchmarked Terra Black against four frontier models.
It ranked first.
100 real Indian title-verification matters. Five AI models. Terra Black scored 0.891 overall and ranked #1, ahead of every frontier model tested.
- macro score
- 0.891
- of 5 systems
- #1
- real matters
- 100
LANDEED PROPERTY INTELLIGENCE BENCH · 100 REAL INDIAN PROPERTY MATTERS
Title verification isn't just a reading test.
It takes judgment.
A title-verification report must do more than extract facts; it must connect the evidence, interpret its implications and reach a defensible conclusion. So we benchmarked 100 real Indian matters against one fixed legal rubric, applied the same way to every model.
100
real Indian title matters scored
500+
real source documents
8
legal dimensions graded
~3–4 min
per full title report
THE LEADERBOARD
Five systems. One benchmark.
Terra Black ranked first.
Across 100 title matters, Terra Black scored 0.891, the highest macro score among the five models tested.
#1 OF 5 SYSTEMS
0.891
macro score: mean criterion pass-rate across 100 title matters
Property cases passed in full: every scored criterion satisfied
- 26Terra Blackof 100
- 14Sonnet-5of 100
- 9Opus-4.8of 100
- 8Gemini-3.6of 100
- 1GPT-5.6of 100
The average scores are close. The all-pass test is less forgiving.
WHERE TERRA LEADS
Frontier models read well.
Terra Black goes further.
The field is strong at extracting facts from property documents. Terra Black's advantage appears where the work requires applying Indian property law and reaching a judgment, while staying competitive across the rest of the title-verification workflow.
Pass rate across eight legal dimensions · 100 real Indian title cases
Pass rate across eight legal dimensions · 100 real Indian title property cases.
| DIMENSION | Terra Black | Sonnet-5 | Gemini-3.6 | Opus-4.8 | GPT-5.6 |
|---|---|---|---|---|---|
| Statutory & tenure | 0.888 | 0.757 | 0.846 | 0.771 | 0.682 |
| Mortgage & security | 0.797 | 0.716 | 0.691 | 0.703 | 0.686 |
| Chain of title | 0.888 | 0.880 | 0.888 | 0.798 | 0.810 |
| Documents scrutinized | 0.938 | 0.938 | 0.926 | 0.907 | 0.895 |
| Title opinion & recs | 0.880 | 0.770 | 0.950 | 0.760 | 0.320 |
| Encumbrances & liens | 0.884 | 0.919 | 0.849 | 0.860 | 0.907 |
| Parties | 0.924 | 0.931 | 0.768 | 0.920 | 0.720 |
| Property identification | 0.893 | 0.917 | 0.938 | 0.930 | 0.909 |
Highest score in each dimension shown in bold. Report structure scored 1.000 for all five and is omitted as a non-differentiator.
Highest score in each dimension marked with a dot. Report structure scored 1.000 for all five and is omitted as a non-differentiator.
#1 OF ALL 5 SYSTEMS
0.888
Statutory & tenure standing
Freehold, RERA, tenure and conversion, joint-family interests and minors' shares, and what they mean for the title.
AHEAD OF THE OVERALL LEADER
0.797
Mortgage & security opinion
Assessing whether the property can support a valid and enforceable mortgage.
CASE BY CASE
The lead isn't just in the average.
It holds case by case.
Across 100 title matters, Terra Black matched or beat each model on a majority of cases and scored higher on average in every head-to-head comparison.
CASE-LEVEL OUTCOMES · 100 TITLE REVIEWS
| Terra higher | Equal score | Opponent higher | mean delta per case → | |
|---|---|---|---|---|
| Terra vs. GPT-5.6 | 87% | 13% | +0.122 | |
| Terra vs. Opus-4.8 | 57% | 19% | 24% | +0.048 |
| Terra vs. Gemini-3.6 | 73% | 27% | +0.039 | |
| Terra vs. Sonnet-5 | 51% | 21% | 28% | +0.029 |
SPEED & QUALITY
The highest score. In about four minutes.
Terra Black produces a complete title-verification report in about four minutes in production, while posting the benchmark's highest score at 0.891. High-quality property intelligence doesn't have to mean a long wait.
BENCHMARK QUALITY × OBSERVED COMPLETION TIME
top-right is the sweet spot
- Terra Black
- Sonnet-5
- Gemini-3.6
- Opus-4.8
- GPT-5.6
Timings are approximate, derived from completion timestamps in benchmark runs.
ROBUSTNESS & SCALE
Real-world property matters, at real-world scale.
The benchmark used real title-verification matters with large, complex document sets. On one of its hardest matters, Terra Black scored 0.833, the highest of the systems tested.
Case size across the benchmark
megabytes
- 9 MB median
- 110 MB largest
- 129 MB hardest
File size describes the workload, not legal difficulty.
The single hardest case (~129 MB)
score, higher is better
| Terra Black | 0.833 | |
|---|---|---|
| Opus-4.8 | 0.722 | |
| Gemini-3.6 | 0.667 | |
| GPT-5.6 | 0.556 |
On the largest case by file size (~110 MB), a frontier model led and Terra Black trailed at 0.733.
METHODOLOGY & DISCLOSURE
One benchmark. Five harnesses. Same standard.
All five models were evaluated on the same 100 title matters, against the same fixed legal rubric, using the same criterion-level scoring process.
ONE FIXED RUBRIC
Every title matter was scored criterion by criterion against the same legal rubric.
CRITERION-LEVEL SCORING
Each requirement was scored independently, then aggregated into matter-, dimension- and benchmark-level scores.
SEPARATE EVALUATOR MODEL
A separate model evaluated every output against the same rubric at temperature 0.
ONE OUTPUT PER TITLE MATTER
One evaluated output per system, per title matter. No best-of-N selection.
Internal benchmark.
See what Terra does with your Property Case.
We benchmarked Terra Black on 100 real title-verification matters. Now try it on one your team knows well and compare the evidence it finds, the analysis it produces, and the opinion it reaches.
Test Terra on a property case →(opens in a new tab)
