VGT AI Benchmark / Observatory
Benchmark explorer
Search and inspect published benchmark evidence with functional server-rendered deep links.
VGT AI Benchmark / Observatory
Model profile
GPT-Image 2.5 Flare by OpenAI
OPENAI
GPT-Image 2.5 Flare
2.5-flare · api
- Verification
- SELF_REPORTED
- Context
- 32.768
- Parameters
- —
Capability profile
Latest completed category scores.
Declared capabilities
text image_input image_output
Run history
Benchmark version, cost evidence and latency remain attached to each run rather than averaged into a vague model label.
Efficiency relationships
Quality is shown against measured cost and latency. Lower x-values are more efficient; the accessible table carries the same evidence.
Calibration
Confidence should track observed correctness. A diagonal indicates ideal calibration.
VGT AI Benchmark / Observatory
Explorer leaderboard
Hard benchmark evidence with explicit version and verification context.
| Model | Provider | Category | Score | Verification | Benchmark |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | coding | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 | |
| Gemini 3.8 Flash | mathematics | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 | |
| Gemini 3.8 Flash | reasoning | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 |