Elektric
Docs

Benchmark

See which models fit your AI requests best on quality and cost.

#What you learn

Benchmark shows how your current model setup compares on expected quality and cost, which models may fit your AI requests better, and where Grid may help.

#How it works

  1. Connect: start a Benchmark and connect your app.
  2. Collect AI requests: keep using your app normally while Elektric collects representative requests.
  3. Understand the requests: Elektric groups the requests by what they are asking the model to do and how difficult they are.
  4. Compare models: Elektric compares your AI requests with its reviewed evidence on model quality, cost, speed, and reliability.
  5. See the results: compare your current setup with models that may offer better quality, lower cost, or both.

Benchmark does not run every AI request through every model. It uses Elektric’s reviewed model evidence to make the comparison.

#Run a Benchmark

  1. Start a Benchmark in the Elektric dashboard.
  2. Follow the connection steps and add the supplied Observe token to your app.
  3. Keep using your app normally while Elektric collects AI requests.
  4. Return to Benchmark when collection is complete to analyze the results.

#What the results show

See how your current model setup compares on expected quality and projected provider cost. Benchmark highlights models that may offer better quality, lower cost, or both, and shows where Grid may be a better fit.

#What Benchmark captures

Benchmark captures the final user prompt and request metadata for AI requests it observes. It does not capture assistant responses, system or developer instructions, earlier conversation turns, tool payloads, or attachment contents.

Captured prompt data is retained for up to 14 days. See the capture and privacy guide for details.

#Limits

  • Benchmark collects up to 1,000 AI requests.
  • Collection runs for up to 7 days after the first captured request.
  • At least 50 AI requests are needed for a useful result.

#Troubleshooting

If the request count is not increasing, check that the current Benchmark token is installed correctly and that your AI client is connected to Observe. Use Benchmark Logs to see accepted, rejected, and failed requests.