Can an agent use your product?

A real coding agent reads your docs and tries one onboarding task in a fresh sandbox. An independent judge verifies the evidence. You get a verdict and, on graded runs, an agent experience score.

Sorted by experience score. Experience measures friction in the tested task. Completion says whether it succeeded.

#ProductScore
Loading benchmark results...
01Open-MeteoWeather API100% passed
Open-Meteo agent profile: Install S, Auth S, Execute S, Docs S, Current S100
02context.devSearch100% passed
context.dev agent profile: Install S, Auth S, Execute S, Docs S, Current S100
03FirecrawlSearch100% passed
Firecrawl agent profile: Install S, Auth S, Execute S, Docs S, Current S100
04ParallelSearch75% passed · 1 blocked
Parallel agent profile: Install S, Auth S, Execute S, Docs S, Current S100
05tinyfish.aiBrowser Automation67% passed · 1 blocked
tinyfish.ai agent profile: Install S, Auth S, Execute S, Docs S, Current S100
06ExaSearch88% passed · 1 blocked
Exa agent profile: Install S, Auth A, Execute A, Docs S, Current S82
07struct.aiAI production monitoring0% passed · 1 blocked
struct.ai agent profile: Install D, Auth C, Execute D, Docs S, Current A26
08apollo.ioSales intelligence0% passed · 2 blockedNot graded
09AI SDKAI SDK0% passed · 1 blockedNot graded
10StripePayments0% passed · 1 blockedNot graded
11tinyfish.comBrowser Automation0% passed · 1 blockedNot graded
11 products
Page 1