VGT AI Benchmark / Observatory
Benchmark explorer
Search and inspect published benchmark evidence with functional server-rendered deep links.
VGT AI Benchmark / Observatory
Model profile
Gemini 3.8 Flash by Google
Gemini 3.8 Flash
3.8-flash · api
- Verification
- VERIFIED
- Context
- 2.097.152
- Parameters
- —
Capability profile
Latest completed category scores.
Declared capabilities
text code tool_use long_context structured_output image_input audio_input video_input
| Category | Score | Benchmark | Version |
|---|---|---|---|
| coding | 100,00 % | VGT AI Benchmark Core Suite 2026 | 1.0.0 |
| mathematics | 100,00 % | VGT AI Benchmark Core Suite 2026 | 1.0.0 |
| reasoning | 100,00 % | VGT AI Benchmark Core Suite 2026 | 1.0.0 |
Run history
Benchmark version, cost evidence and latency remain attached to each run rather than averaged into a vague model label.
| Benchmark | Finished | Categories | Cost evidence | Latency evidence |
|---|---|---|---|---|
| VGT AI Benchmark Core Suite 2026 1.0.0 | 2026-09-27 | 3 | {"source":"unknown"} | 33 ms |
Efficiency relationships
Quality is shown against measured cost and latency. Lower x-values are more efficient; the accessible table carries the same evidence.
| Benchmark | Average score | Cost | Latency |
|---|---|---|---|
| VGT AI Benchmark Core Suite 2026 1.0.0 | 100,00 % | — | 33 ms |
Calibration
Confidence should track observed correctness. A diagonal indicates ideal calibration.
| Confidence | Observed accuracy |
|---|---|
| 95,00 % | 100,00 % |
VGT AI Benchmark / Observatory
Explorer leaderboard
Hard benchmark evidence with explicit version and verification context.
| Model | Provider | Category | Score | Verification | Benchmark |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | coding | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 | |
| Gemini 3.8 Flash | mathematics | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 | |
| Gemini 3.8 Flash | reasoning | 100,00 % | VERIFIED | BEN_vgt_core_suite_2026_v1 1.0.0 |