Internal Ukrainian-law benchmark · 6 August 2026
The same model achieved a higher legal-quality score inside PRISM
On the same 40-test Ukrainian-law benchmark, GPT 5.6 Solo MAX scored 154/160 inside PRISM4 versus 147/160 standalone, while both DeepSeek configurations also improved.
Headline comparison
GPT 5.6 Solo MAX · Inside PRISM4
154/160 · 96.3%
3.85/4 · 36/40 perfect responses
GPT 5.6 Solo MAX · standalone site
147/160 · 91.9%
3.68/4 · 29/40 perfect responses
40-response quality strips
Inside PRISM4
36/40
Standalone
29/40
+7 points
+7 perfect responses
PRISM result on the same 40-test set. The model alone is not the complete product.
Internal Ukrainian-law benchmark, 6 August 2026. Forty tests; 160 maximum points.
Base mode versus orchestration mode
Orchestration improves results across different base models
DeepSeek V4 Pro MAX
+5 · +3.1 pp
Base 144/160 · 90.0% → orchestration 149/160 · 93.1%
DeepSeek V4 Flash MAX
+9 · +5.6 pp
Base 127/160 · 79.4% → orchestration 136/160 · 85.0%
90.0%
144/160
Base
93.1%
149/160
Orchestration
79.4%
127/160
Base
85.0%
136/160
Orchestration
Method
One test set. Two operating modes.
Forty Ukrainian-law tasks were scored on a four-point scale. Each base model was compared with its PRISM4-orchestrated configuration, for a maximum of 160 points per run.
01 · Tests
40
02 · Points each
4
03 · Maximum
160
04 · Run date
6 Aug 2026
Complete scorecard
GPT 5.6 Solo MAX
- Standalone
- 147/16091.9%
- Inside PRISM4
- 154/16096.3%
- Gain
- +7+4.4 pp
DeepSeek V4 Pro MAX
- Standalone
- 144/16090.0%
- Inside PRISM4
- 149/16093.1%
- Gain
- +5+3.1 pp
DeepSeek V4 Flash MAX
- Standalone
- 127/16079.4%
- Inside PRISM4
- 136/16085.0%
- Gain
- +9+5.6 pp
Every tested configuration scored higher with PRISM4 orchestration.
