Model benchmark
E-learning generation quality across AI models — methodology v1. Every cell runs the full production pipeline with the candidate model pinned (no fallback).
Super-admin only
Your account does not have the super_admin role — the endpoints will refuse. Ask the platform owner.
New run
One Requesty model id per line — exactly these are benchmarked, no fallback
Medians need ≥3
Results
No completed cells yet.
Methodology v1
Each cell pins the candidate model as the ONLY content model (failures are recorded, never rerouted) and generates a text-only lesson through the production pipeline, self-repair passes included. Judges are fixed: anthropic/claude-opus-5, google/gemini-3.1-pro-preview — two families, averaged.
Briefs
- Gifts & invitations under an anti-bribery policy — A lesson for employees of a mid-size medical device company on handling gifts, hospitality and invitations from business partners under a typical anti-bribery policy: value thresholds, approval routes, cultural sensitivities, and how to decline gracefully without damaging the relationship. Outcome: Decide for a concrete gift or invitation whether to accept, decline, or seek approval, and justify the decision by policy criteria
- Structured patient handover with SBAR — A lesson for nursing staff on structured shift handover using SBAR (Situation, Background, Assessment, Recommendation): why unstructured handovers lose critical information, how to build each SBAR component from a real patient case, and how to receive a handover actively. Outcome: Formulate a complete SBAR handover for a given patient case and identify the missing component in a flawed handover
- The 3-2-1 backup rule in practice — A lesson for small-business IT administrators on designing a backup strategy around the 3-2-1 rule: three copies, two media types, one off-site. Covers realistic failure scenarios (ransomware, hardware failure, accidental deletion), restore testing, and common false-security traps like sync services mistaken for backups. Outcome: Design a 3-2-1-compliant backup plan for a described company setup and spot the gap in a non-compliant plan
- Giving corrective feedback that lands — A lesson for first-time team leads on delivering corrective feedback: separating observation from interpretation, anchoring feedback in specific behaviour and impact, choosing the right moment and setting, and handling defensive reactions without escalating. Outcome: Rewrite a judgmental feedback statement into a behaviour-impact formulation and choose an appropriate response to a defensive reaction
Judge rubric (verbatim)
- Outcome alignment: Does every section serve the stated performance outcome? Is the outcome actually practised and measured, not just mentioned? Penalize content that is on-topic but does not move the learner toward the outcome.
- Instructional design: Is there a deliberate arc (activate prior knowledge → teach → practise → assess)? Are concepts scaffolded before they are tested? Are examples concrete and did the lesson use realistic cases rather than abstract description?
- Assessment quality: Do quiz questions test application and judgment rather than recall of the preceding paragraph? Are distractors plausible? Is the correct answer NOT guessable from wording alone? Are explanations instructive?
- Writing quality: Professional, concrete prose a subject-matter colleague would sign off. Penalize: filler, repetition, hedging, motivational padding, formulaic AI constructions (parallel negation fragments, rhetorical triplets, "isn't just X — it's Y" reveals), and invented statistics.
- Factual soundness: Is the domain content correct and current? Flag concrete errors, invented specifics, or advice a practitioner would consider wrong or risky. Judge only what the lesson asserts.