Back to news

Internal Ukrainian-law benchmark · 6 August 2026

The same model achieved a higher legal-quality score inside PRISM

On the same 40-test Ukrainian-law benchmark, GPT 5.6 Solo MAX scored 154/160 inside PRISM4 versus 147/160 standalone, while both DeepSeek configurations also improved.

PRISM4 Research

Headline comparison

GPT 5.6 Solo MAX · Inside PRISM4

154/160 · 96.3%

3.85/4 · 36/40 perfect responses

GPT 5.6 Solo MAX · standalone site

147/160 · 91.9%

3.68/4 · 29/40 perfect responses

Score axis · 120–160
Inside PRISM496.3%
Standalone91.9%
120130140150160

40-response quality strips

Inside PRISM4

36/40

Standalone

29/40

+7 points
+7 perfect responses

PRISM result on the same 40-test set. The model alone is not the complete product.

Internal Ukrainian-law benchmark, 6 August 2026. Forty tests; 160 maximum points.

Base mode versus orchestration mode

Orchestration improves results across different base models

DeepSeek V4 Pro MAX

+5 · +3.1 pp

Base 144/160 · 90.0% → orchestration 149/160 · 93.1%

DeepSeek V4 Flash MAX

+9 · +5.6 pp

Base 127/160 · 79.4% → orchestration 136/160 · 85.0%

90.0%

144/160

90.0%

Base

93.1%

149/160

93.1%

Orchestration

DeepSeek V4 Pro MAX

79.4%

127/160

79.4%

Base

85.0%

136/160

85.0%

Orchestration

DeepSeek V4 Flash MAX

Method

One test set. Two operating modes.

Forty Ukrainian-law tasks were scored on a four-point scale. Each base model was compared with its PRISM4-orchestrated configuration, for a maximum of 160 points per run.

  1. 01 · Tests

    40

  2. 02 · Points each

    4

  3. 03 · Maximum

    160

  4. 04 · Run date

    6 Aug 2026

Complete scorecard

GPT 5.6 Solo MAX

Standalone
147/16091.9%
Inside PRISM4
154/16096.3%
Gain
+7+4.4 pp

DeepSeek V4 Pro MAX

Standalone
144/16090.0%
Inside PRISM4
149/16093.1%
Gain
+5+3.1 pp

DeepSeek V4 Flash MAX

Standalone
127/16079.4%
Inside PRISM4
136/16085.0%
Gain
+9+5.6 pp

Every tested configuration scored higher with PRISM4 orchestration.