October 14, 2026
Second full field test with eight STQ maintenance crew members, testing the revised prompt architecture with context-building clarifying questions and the redesigned touch interface.
Prompt architecture. The clarifying-question logic worked. When a crew member said “the pressure is off,” the assistant responded with a two-option prompt: “Which system — hydraulic or pneumatic?” Crew members found this natural and answered quickly. Response relevance ratings from crew members (1–5 scale, verbal) improved from an average of 2.8 in the first test to 4.1 in this test.
Touch interface. Glove-compatible mode held up in both open-air and enclosed-bay conditions. Tap targets sized for nitrile gloves worked; crew members stopped fighting the interface and started using it.
New failure mode. Two crew members asked the assistant follow-up questions after a procedure was complete — essentially using it as an after-action reference. The assistant handled these queries correctly in isolation, but without session continuity, it had no context from the earlier procedure. Crew members had to re-establish context each time, which was frustrating. Session memory is now on the development roadmap.
Disagreement logging. We deployed the crew-AI disagreement logging protocol. In this session, crew members flagged four instances where assistant output conflicted with their practice. Three were cases where the assistant was technically correct but the crew’s adapted practice was safer given local equipment variation. One was a model error. All four are now tagged for review.
The system is beginning to feel like a tool rather than a demo. That is the threshold we were trying to cross.