TL;DR: A test of Claude, Gemini, and o3 on a tree-based combobox found that LLMs can scaffold compound component APIs quickly, according to WorkOS, but they still struggle with nested behaviour, keyboard support, screen-reader semantics, and state coordination in complex UI. That makes context, tests, and manual review the real guardrails, not prompt length alone.
At a glance
What this is: WorkOS tested LLMs on a tree-based combobox and found that they could generate a strong scaffold, but the component still broke on accessibility, nested interaction, and state coordination.
Why it matters: For IAM and developer-platform teams, the lesson is that AI-assisted UI generation speeds scaffolding but does not replace behavioural testing, accessibility validation, or human review for complex interactive flows.
Context
A tree-based combobox is a search-driven, nested selection control where parent nodes can expand to reveal child items. In this case, the governance gap was not whether an LLM could write React code, but whether it could preserve complex interaction rules across composition, keyboard use, and accessibility semantics.
The article matters to identity and access teams because similar failures appear whenever interfaces must enforce precise state, selection, and navigation behaviour. In human IAM and admin tooling, small UI mistakes can become workflow blockers, accessibility defects, or control gaps if teams rely on generated code without verifying how it behaves under real use.
Key questions
Q: How should teams use LLMs safely for complex UI components?
A: Use LLMs for scaffolding, boilerplate, and pattern completion, but require tests and human review for behaviour, accessibility, and state coordination. The safest workflow is to treat the model’s output as a draft that must pass keyboard, focus, and screen-reader checks before merge. For identity-adjacent UI, the acceptance bar should be stricter, not looser.
Q: What breaks when a generated component looks correct but the interaction model is wrong?
A: The usual failure points are nested state, focus management, keyboard navigation, and accessibility semantics. A component can render cleanly and still fail users if its event handling, hierarchy, or screen-reader roles do not match the intended behaviour.
Q: How do security and platform teams know when to stop prompting and rewrite code manually?
A: Stop when repeated prompts are producing regressions, when stateful behaviour keeps drifting, or when testing shows that the tool understands the shape of the component but not the logic behind it. At that point, manual implementation is usually faster than continued correction.
Q: Why do complex, nested controls need more than code generation to be reliable?
A: Because nested controls combine search, hierarchy, selection, and accessibility into one state machine. Generating the surface API is relatively easy, but maintaining correct behaviour across all those layers requires design intent, tests, and careful review.
Technical breakdown
Why compound component APIs are hard for LLMs
Compound components depend on a consistent relationship between parent, child, and context providers. A model can imitate the shape of a Radix-style API, but that does not guarantee the nested state model actually works. In this case, the generated output often looked structurally plausible while missing the deeper constraints that make composition usable, such as where state lives, how items register, and how children inherit behaviour. That gap matters because UI scaffolding and UI correctness are different problems. The model can reproduce familiar patterns, yet still fail at the coordination logic that makes the pattern functional.
Practical implication: Treat generated component structure as a draft, not as proof that the interaction model is valid.
Accessibility semantics and keyboard flow in nested widgets
Accessibility in a tree combobox is not just a matter of adding ARIA labels. Nested widgets need coherent roles, predictable focus movement, and keyboard behaviour that matches user expectations. The article shows how a design can look right visually while still breaking for screen readers or keyboard users when collapsible state, selection, and navigation collide. That is a common failure mode in AI-assisted UI work: the output may render, but assistive technology support remains fragile because semantics and interaction state were not designed together. Accessibility has to be tested as behaviour, not inferred from markup patterns.
Practical implication: Validate nested widgets with screen readers and keyboard-only testing before the component reaches production.
Why search and filtering logic broke on tree data
Tree-shaped data changes how filtering must work. A child match may require the parent to stay visible and expanded, while a parent match may hide descendants until expansion. The article’s failures show how LLM-generated code often searches the wrong values, filters the wrong layer, or loses the hierarchy needed to preserve context. That is especially important in component libraries because filtering is not a cosmetic feature, it is part of the control’s state machine. Once filtering and expansion interact, the implementation needs explicit rules, not just a generic search routine copied from flatter patterns.
Practical implication: Define filtering rules for hierarchical data before asking an LLM to implement the component.
NHI Mgmt Group analysis
LLM-assisted UI generation is a scaffolding accelerator, not a behaviour guarantee. WorkOS’s example shows that the first pass can be structurally close while still failing on the hard parts: nested state, keyboard interaction, and accessibility semantics. That distinction matters across identity and admin tooling, where the component may look complete but still fail under real user behaviour. The practitioner conclusion is that generated code should be judged on interaction fidelity, not visual completeness.
Complex UI surfaces the limits of pattern matching. LLMs are strongest where the target resembles prior code in the training distribution, and weakest where the UI introduces uncommon state coordination. A tree combobox combines hierarchical selection, expansion, filtering, and assistive-technology behaviour in one control, which makes it a poor candidate for blind generation. The broader lesson is that novelty in interaction design lowers the odds that prompt-only iteration will converge quickly.
Context beats prompt length when behaviour is nuanced. The article reinforces a practical truth: tests, file-level comments, and embedded design-system context create a more durable specification than chat history alone. That is especially relevant for teams building internal admin surfaces, where accessibility and state consistency are part of the control plane. The practitioner takeaway is to make the codebase itself carry the intent, not the conversation.
Prompting cannot substitute for verification when semantics and state interact. The failure was not just that the component was incomplete, but that further prompting sometimes made it worse by drifting away from the intended composition model. This is the key governance insight for AI-assisted development: the review process has to prove behaviour, not admire output. The practitioner conclusion is to move validation earlier, because late-stage prompting is a weak substitute for test-driven correction.
Named concept: behavioural gap debt. This article shows the cost of accepting a generated scaffold before the edge cases are proven. In nested UI, that debt appears as accessibility regressions, broken focus, and state collisions that only emerge during manual testing. The practitioner conclusion is to treat behavioural gap debt as a release risk, not a coding inconvenience.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: AI Coding Agents Security Guide
What this signals
Behavioural gap debt: AI-generated UI often accumulates hidden defects after the scaffold is in place, especially when focus, hierarchy, and selection rules interact. Teams should expect the first implementation pass to look complete while still requiring test-driven correction before release.
The practical boundary is simple: if a control must satisfy keyboard users and screen-reader users, the code generator is only the starting point. For identity and admin surfaces, the safest workflow is to specify behaviour in tests and comments before the LLM writes the first pass.
For practitioners
- Start with behaviour tests Codify keyboard paths, expansion rules, and screen-reader expectations before using generated code so the model has an explicit contract to satisfy.
- Embed component context in files Use comments, local examples, and design-system patterns inside the codebase so intent survives after chat context is lost.
- Review accessibility as runtime behaviour Test nested widgets with screen readers and keyboard-only flows, because correct-looking markup can still fail users in practice.
- Limit prompt-only iteration for stateful UI When selection, filtering, and focus interact, switch from repeated prompting to manual implementation or refactoring once regressions appear.
Key takeaways
- LLM-assisted UI generation can accelerate scaffolding, but it does not reliably preserve the behaviour of complex nested components.
- Accessibility, keyboard support, and state coordination are the points most likely to fail when a generated control moves beyond familiar patterns.
- The most effective guardrails are explicit tests, durable in-file context, and early manual review when the interaction model becomes nuanced.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while OWASP ASVS and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article concerns LLM-assisted code generation that misapplies component patterns and behaviour. |
| Recommendation — Constrain LLM-generated UI changes to tested component boundaries and verify tool output against expected behaviour. | ||
| OWASP ASVS | V3 — Web Frontend Security | The article’s core failure mode is incorrect frontend behaviour and interaction handling. |
| Recommendation — Verify frontend components for keyboard flow, semantics, and state consistency before release. | ||
| NIST CSF 2.0 | PR.AT-01 — All users are provided awareness and training | The article highlights the need for developer training on LLM limitations and validation discipline. |
| Recommendation — Train developers to validate generated code instead of accepting AI output as production-ready. | ||
Key terms
- Compound Component: A compound component is a UI pattern built from multiple coordinated parts that share state and behaviour. In practice, it lets teams compose interfaces from reusable subcomponents, but it also demands tight control over focus, events, and accessibility so the pieces behave as one coherent control.
- Tree combobox: A searchable selection control that presents options in a hierarchy, with parent items that can expand to reveal child items. It is harder to implement than a flat dropdown because filtering, navigation, and accessibility must preserve the tree structure.
- Accessibility semantics: The roles, attributes, and interaction patterns that let assistive technologies interpret a user interface correctly. In complex controls, semantics must match real behaviour, not just the visible layout, or screen readers and keyboard users will experience broken workflows.
- Behavioral Verification: Behavioral verification uses patterns such as normal device use, login habits, and contextual signals to decide whether an access attempt looks legitimate. It strengthens authentication by comparing the current attempt with expected user behaviour, helping detect account takeover attempts that reuse stolen credentials.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org