[ Models ]
Origin for speed. NextGen for depth.
Two models. One assistant. One API. Compare them, then try in Quark.
One prompt. Two paths.
Quark routes every request to the right model. Pick a scenario and watch the decision.
Prompt
“Answer a support question in real time.”
Quark router · weighing short-form · latency-critical · high volume
Origin
NextGen
Routed to
Origin
128K · ~18ms · speed path
Speed path. Flash-class answer in ~18ms.
Same routing runs live inside Quark
Compare
- MMLU
- 81.4%
- GPQA Diamond
- 48.1%
- HumanEval
- 78.9%
- MATH
- 74.2%
- Context window
- 128K
- Latency
- ~18ms