How to research and design for Dynamic systems
The research playbook for AI products - that are never the same twice
I laid out the methodology shift AI forces on design research: live models over scripted mockups, red-teaming over happy paths, longitudinal logs over lab sessions, rulesets over personas. This is the operational sequel: how to actually install those four shifts in a working team, in what order, and where the installation fails.
AI Key Takeaways
The four shifts of dynamic systems research are a system, but teams shouldn’t adopt them as one. There is a correct order, and it starts with the cheapest change: the boundary-testing brief.
Legacy methods are calibrated to a product that stays fixed after shipping.
The most common adoption failure is hybrid theater: teams keep the scripted Figma prototype and add one “AI question” at the end, which produces the old data with new vocabulary.
The second failure is tooling panic. Live-model sandboxes require an API playground and a written system prompt, and that’s all they require. Waiting for a platform purchase is a stall.
Full adoption inverts the researcher’s role: from validating what the team built to mapping how the system behaves.
Every deliverable research produces has the same reader: a human in a room. The persona, the journey map, the readout with its recommendations slide, written to be read, nodded at, and remembered imperfectly.
AI products change the addressee. The most useful thing research can hand over now is something the system itself can consume: rules the model runs, failure cases the evals replay, constraints engineering can query later. Everything that follows: the four shifts, the order they install in, the ways the installation fails, is downstream of that one switch:
the reader is the model, not the room.
But, the argument for dynamic systems research hasn’t landed.
A year of conversations taught me that practitioners keep defaulting to their static-system audits and their fixed analysis frameworks. Researchers, designers, developers, they don’t yet see what a probabilistic product asks of research, or why methods built for fixed screens can’t answer it.
Part of this is comfort. These are people who are good at what they do, and expertise makes it hard to notice when the object of that expertise has changed underneath you. Part of it is awareness. Trust that forms over weeks, behavior that drifts at the boundaries: those questions never appear on a static audit, so nobody misses the answers.
This piece has two jobs. Make the case, then show how a working research team learns to evaluate behavior, in an order that survives a real roadmap.
One reassurance first, because it defuses the resistance that kills most rollouts. Your methods aren’t wrong. They’re calibrated to a product that no longer exists. Usability scripts, concept tests, personas: each was a correct answer to a world where you design a thing, ship the thing, and the thing stays.
Step one: change the brief
Start with the cheapest shift, the one that requires no new tools, no budget line, and no permission: replace the happy-path task script with a boundary-testing brief. Instead of “ask the assistant to help you find a flight,” the instruction becomes “try to confuse it, contradict yourself, change your mind halfway.” Same lab, same participants, same hour.
This goes first for a political reason as much as a methodological one. The first boundary-testing session produces findings the team has never seen: the model stalling, the interface hiding uncertainty, the user failing to notice a hallucination for four turns. Those findings sell the rest of the transition better than any argument.
A research leader who starts with tooling has to justify cost before showing value. A leader who starts with the brief shows value in week one.
Step two: put a live model in the prototype
The second shift replaces hardcoded responses with a live model driven by authored system prompts.
The moment a response is scripted, the core variable of AI design (uncertainty) has been removed, and the study is ranking the researcher’s own copywriting.
The fix costs an afternoon: write two or three system prompts encoding the tonal or behavioral variants you’d otherwise have scripted, put participants in front of an API playground, and let their questions vary the way real questions vary.
The failure mode here is tooling panic: teams stall for a quarter waiting for a research platform to add “AI testing features.” Nothing about this step requires procurement. A playground, a system prompt, and a note-taker constitute the entire stack. The team that runs its first live-model study on borrowed API credits is two quarters ahead of the team waiting for a vendor.
Step three: extend the study past the session
The third shift is where the calendar and the budget actually move: from single-session evaluation to multi-week prompt-stream analysis. Trust in a probabilistic system is cumulative. It forms and erodes over repeated, unmonitored use, and no 45-minute session can observe it. Participants use a live prototype inside their real workflows for two to three weeks, and the conversation logs become the primary data: how many turns to an acceptable output, where rephrasing cascades signal friction, the exact moment a user abandons the tool for their old workflow.
This is the step that requires stakeholder management, because it changes the shape of research deliverables from a readout after a study week to a trust curve after a month.
First impressions are the one thing we already know how to measure, and the product’s fate is decided in week three, long after minute thirty.
Step four: change what you hand engineering
The last shift converts the output. Static persona documents give a model nothing to operate on; behavioral rulesets give it constraints: when the user shows crisis signals, shorten outputs and confirm explicitly; when the user is exploring, tolerate ambiguity and offer branches. Written well, the ruleset doubles as an evaluation dataset, which means research stops delivering inspiration and starts delivering testable specification.
This shift goes last because it depends on the other three: you can’t write behavioral rules for contexts you’ve never observed, and you can’t observe them in a scripted session.
The lifecycle changes with the format.
A readout is terminal: it gets presented, discussed, and filed, and nothing downstream consumes it.
A ruleset is generative: it feeds the next round of probing, and once it passes its own evals, it ships. The old deliverable ended with a walkthrough. The new one doesn’t have an ending: it has a next phase.
The two failure modes to watch
Adoption fails in two recognizable ways:
The first is hybrid theater: the team keeps the scripted prototype, keeps the happy-path script, and appends an “AI question” to the discussion guide. The deliverables look updated and the data is unchanged. The tell is a readout that still reports task completion rates for a system whose tasks are open-ended.
The second is skipping to the tooling. A team buys a logging platform before it has changed a single brief, and the new dashboard fills with data nobody has a question for. The sequence exists because each step generates the demand for the next: adversarial sessions surface behaviors worth testing live, live testing surfaces patterns worth tracking over weeks, and longitudinal patterns surface the rules worth handing to engineering.
The researcher’s job stops being validation of what the team already built and becomes cartography of how the system actually behaves. That second job is the one the next decade pays for.
NA: AI-assisted tools were used for transcription, reference formatting, and language editing. All intellectual content and conclusions remain solely the author’s.







