Your users are already telling you what is broken.

VibeCheck listens to the shape of a journey, never its content. Jev reads where people struggle, and the strongest signals become short studies with real people.

Taking part in a study? →

Open the demo
A visitor wants an image of their drawingsemantic log · excalidraw
0:04.1journey_startjourney_id=share_drawing
0:04.1progressprogress_ref=drawing_started
0:07.9action_attemptaction_ref=share_button
0:09.3action_resultaction_ref=share_button result=cancelled
0:13.1action_attemptaction_ref=save_to_file
0:15.0action_resultaction_ref=save_to_file result=cancelled
0:17.4help_requesttarget_ref=help_dialog
0:21.0navigationroute_template=/menu/main
0:22.6action_resultaction_ref=export_image result=success
0:22.6completionprogress_ref=image_exported

What Jev reads into it · evidence sufficient · share mistaken for export

Friction observed95%
Wrong path taken97%
Recovered after help97%
Research warranted81%

Above the gate → research candidate → study

From a recorded demo run, model jev-1.13.0. An example, not a benchmark.

help_request

0:17.4

Window in the last 90 s?

yesno

Jev evaluates the window

friction 95% · wrong path 97% · candidate

keeps listening

no model call

Triggers are code, not the model: a help request, a repeated failure, a navigation loop or the journey ending. Jev only reads what the trigger cut, bounded by budget and cooldown.

Signals, not findings. Until real people confirm them.

Nothing on screen is simulated: every number is a stored event, a real model answer or a verified recording. Model estimates are labeled as estimates, sample data as sample.

Open the demo