Pathwright AI Research & Data Science
Role: AI research and data science intern
Tools: Python, notebooks, structured evaluation tables, stakeholder summaries, PathMX
Event: Grok Conf ’26 (Pathwright sponsor; PathMX conference web experience)
Summary
At Pathwright, I am helping explore how AI-assisted instructional tools can support more personalized learning. The work sits between data science, product research, and education: define the question clearly, collect comparable evidence, and explain model behavior in a way that helps the team make a decision. I have learnt a lot from my supervisor and embraced the opportunity to work on cutting-edge innovations in the field of AI and education.
That research also showed up live at Grok Conf ’26, where Pathwright sponsored the conference and used the web experience as a place to try early PathMX technology. I helped build out PathMX content for Grok—an early, public stress test of curriculum-as-Markdown outside the usual research notebook.
Problem
AI agents can produce useful instructional support, but their behavior needs to be evaluated consistently. A single good response is not enough; the team needs repeatable ways to compare outputs, spot failure modes, and decide what should change in prompts, product flows, or evaluation criteria. In addition, in order to make coding more engaging and accessible, we need to explore how AI can be used to help students learn to code by creating personalized learning paths.
My Role
- Help translate broad product questions into measurable research tasks.
- Organize model outputs into structured observations.
- Compare results across prompts, user contexts, or evaluation criteria.
- Summarize findings for teammates who need clear decisions instead of raw transcripts.
- Prepare personalized guides for students for each lesson to help them learn to code by themselves.
- Contribute PathMX Sources and supporting work for the Grok Conf ’26 conference experience.
Approach
1. Frame the Evaluation
Start with a narrow question, such as whether an instructional agent identifies the learner's actual blocker or whether it gives generic advice.
2. Capture Comparable Outputs
Collect responses in a format that makes them easy to compare across the same dimensions: accuracy, specificity, helpfulness, tone, and next-step clarity.
3. Tag Failure Modes
Group recurring issues into practical categories, such as vague guidance, missed context, overconfident claims, or weak follow-up questions.
4. Report the Decision
Turn the analysis into a concise recommendation: keep, revise, test again, or investigate further.
Grok Conf ’26
Grok Conference is a design-and-tech gathering Pathwright has long supported. For Grok ’26, the team treated the conference web UX as a live lab for early PathMX: Markdown-first Sources built into a playable experience for attendees, alongside other experiments like NFC badge interactions.
My part was helping author and shape PathMX material for that Grok surface—turning research ideas about durable, playable learning content into something people could open during the conference. Working under time pressure with unfinished tooling made the usual research habits concrete: keep Sources readable, make structure obvious, and notice where the learner (or attendee) gets stuck.
That conference work connects directly to the rest of my internship: evaluate what an agent or curriculum produces, then put the durable artifact somewhere people can actually use it—not only in a chat or a private notebook.
Example Artifact
A useful artifact for this case study would be a one-page evaluation summary:
- Research question.
- Sample size and source.
- Scoring dimensions.
- Top 3 failure modes.
- Recommended product or prompt change.
What This Shows
This project shows applied data thinking in an AI product setting: defining evidence, structuring messy outputs, and communicating what the team should do next. Grok ’26 added a second layer—shipping early PathMX into a real event audience, where clarity and recoverability matter as much as the research claim.
Next Steps
- Add a sanitized screenshot or table from an evaluation run.
- Include one before/after example showing how a prompt or workflow changed.
- Add a short metric summary once there is enough comparable data.
- Link a public Grok / PathMX artifact or screenshot from the conference experience when available.