Six testers: Ben · Emmy · Than · Neung · Tong's mother · Parin. 12 to 14 August 2026.
Five of six independently said they could not tell when it was over. Across every session ever recorded, real users who reached a closing step: zero. Fix order matters: problems 2 and 5 are infrastructure, and until they are fixed, any test of anything else is measuring the wrong thing.
One thing not to flatten: the testers disagreed about whether the tool understands people well, and the disagreement tracks how many problems each of them brought. See the section near the end before repeating "everyone liked the comprehension", which is not true.
Everyone hit it. It is why everyone dropped.
What happens
The conversation never terminates. No closing screen, no completion signal, no statement of what the user walks away with. People talk until they tire, then close the tab.
Ben: "ละตอนจบแบบว่ากุงงๆ แบบไม่รู้ว่าเออจบละหรอ เหมือนมันก็เหมือนจีพีทีอ่ะ คือมันก็ถามต่อไปเรื่อยๆ ซึ่งกุไม่รู้มันมีจุดจบไหม"
Emmy: "ในจว่าตรงนี้คือจบยังง" and "ไม่รู้ว่าจริงจริงต้องพูดจบถึงตรงไหน"
She asked twice, 90 minutes apart.
Neung: "ไม่รู้ว่าต้องหยุดพูดตอนไหน ยังไม่รู้ว่าตอนนี้เราพูดวนไปวนมา มากน้อยขนาดไหนแล้วหรือยัง"
Than: "ตอนสุดท้าย Result มันจะได้ออกมาหน้าตาเป็นแบบไหน ระหว่างทางมันก็เลยไม่ได้รู้ว่าต้องช่วยปรับตรงไหน"
The step counter makes it worse. Ben saw "ขั้นที่ N จาก 4" and expected something at 4: "กุนึกว่าแบบตอนจบ 4 มันจะมีแบบเหมือน End journey sign or sth". Screenshots confirm "ขั้นที่ 4 จาก 4" is still an open chat box asking questions. Than could not work out what a step even counted: "แต่ละขั้นมันหมายถึงจำนวนข้อคำถาม หรืออะไรยังไง".
Fix
Accept when
audit_people.py shows a non-zero count at the closing step.100% of recent real sessions. This is a bug, not a design problem.
What happens
The board asks for six question cards, gets three, and displays the short set anyway. Then it refills continuously, so questions appear to swap, repeat and pile up. It has done this 910 times: asked for six and got three 604 times, got two 155 times, got one 26 times.
| session | times the board came back short |
|---|---|
| 13 Aug 23:35 | 392 |
| Emmy, 13 Aug 15:19 | 135 |
| Neung, 14 Aug 10:43 | 118 |
| Ben, 12 Aug 22:02 | 36 |
Emmy: "มันจะมีบางแวบที่รู้สึกว่า เอ๊ะ เหี้ย เยอะมาก ไม่รู้ต้องพูดอันไหนก่อน" and "บางคําถามมันก็เหมือนแบบจะถามซ้ำนิดหนึ่ง"
Neung: "ตอนนี้รู้สึกว่ามันยังจับปัญหาได้ไม่ค่อยดี มันถามสลับไปสลับกันมา"
Rewriting the questions will not touch this. The generator is not returning enough candidates. See problem 5 for the likely upstream cause.
Neung found a second cause, separate from the bug
When someone has several problems, they tell them interleaved, and the adaptive questions then interleave too. His words: "สมมุติว่าคนเล่าหลายประเด็นอย่างนี้ คนก็จะเล่าสลับกันไป สลับกันมา แล้วคำถามมันก็จะสลับกันไปสลับกันมา". So even with the board fixed, a person with a crowded head gets a shuffled conversation.
His fix is specific, and it is two things, not one:
Both are aimed at the same thing, which he says plainly: "ก็ช่วยทำให้คนโฟกัสได้". Colour is how you see the grouping without reading; the triage card is how you choose. Doing only the grouping and not the colour loses the part that works at a glance.
Fix
Accept when
board_degraded_held events.Caused a confirmed drop-out. The only failure here with a named casualty.
What happens
The Reflection card asks "we understood it this way, right?" and offers a way to say no. Tapping it and explaining leaves the screen unchanged. The user never finds out whether the correction landed or what the tool now thinks.
Tong, on his mother: "เพราะกดไม่ใช่ไปแล้วอ่ะ แล้วอธิบาย มันก็ยังอยู่ตรงนั้น แล้วมันก็ไม่รู้ว่าสุดท้ายเราเข้าใจตรงกับ user จริงๆ หรือเปล่า ก็ทำให้แม่ drop เราออกไป"
Second fault, at the very start of her session. She tapped a question card once at the beginning, told a long story, and then never tapped another one for the rest of the session: "แม่ก็คลิกแล้วก็เล่าไปยาวๆ แล้วหลังจากนั้นน่ะ แม่ไม่ได้คลิกคำถามอะไรต่อเลย". So the panel is discoverable once and then stops reading as something you keep using. Tong's diagnosis: "คนแก่ไม่รู้ว่าสุดท้ายอันนี้เป็น Adaptive Question นะ คุณต้องกดนะ ตอนนี้คนเขาเล่ายาว ๆ ไปเลย". Nothing teaches people that the cards are how you steer.
Third, the wording. Emmy says "ตรงนี้ไม่ตรง พูดแก้" feels like being marked on homework, "มันเหมือนจะเป็นตรวจงานไปนิดนึง", and she often cannot remember what she said anyway.
The three pieces, so the fix targets the right one
Tong names the product as three components, and the failures land on different ones: Open-ended Questions (the panel of cards), Reflection (what appears while you talk, checking understanding), and Summary (a page of cards to tick, then a combined summary). His mother was lost at the panel and left at the Reflection. She never reached either Summary page.
Fix
Accept when
Corrupts the input, not just the mood.
What happens
The screen is emphatic about how much we collect and silent on what happens next and for how long. Thorough about taking, quiet about keeping, reads as a warning.
Than: "ติดหลักๆ privacy💀 สิ่งที่กุกลัวคือหน้าแรกก่อนเริ่มอะ เราจะเก็บข้อมูลคุณนะ แต่ไม่รู้บริษัททำไรต่อ เราเก็บข้อมูลนะ แต่ไม่รู้นานแค่ไหน เหมือนพี่ชายเน้นว่าเอาข้อมูลไปใช้เยอะ จนคนอ่านระแวงใจ"
The consequence is bigger than the complaint. He deliberately gave shallow answers: "เลยไม่ได้ให้ข้อมูลขนาดนั้น เลยอาจไม่เห็นประสิทธิภาพมันจริงๆ". His session is not a valid test of the product. The consent screen is suppressing the data the product runs on.
Note this gets larger, not smaller, if the positioning is "we sell what is inside a person's head at work". Then what the paying organisation can see becomes the central product question, and two extra sentences will not settle it.
Fix
Accept when
Neung's session this morning. Second occurrence: 30 July was the first.
What happens
Every question the tool generates is a call to Google's Gemini. The account has a fixed number
of calls allowed per minute and per day. Past that, Google returns 429 Too Many Requests
instead of an answer. The product runs on a personal account which has hit its ceiling.
In Neung's session this happened 18 times, three retries deep each time, starting at 5 minutes 36 seconds and continuing to the end of his 7-minute session. He carried on talking through it. Five turns were captured after the refusals began and none of them could be processed. The last minute of his session was the tool silently doing nothing.
Separately, the model returns unusable output constantly: 149 400 Bad Request, plus
over 100 truncated-JSON failures ("Unterminated string starting at line 2, line 3, line 8").
That is a known failure shape when thinking output is uncapped, and it is the most likely reason
the question board keeps coming up short.
Fix
Accept when
A product decision, not a UI one. Blocks the UX redraft.
What happens
The tool addresses two different problems and commits to neither: problems in the work (structure, a manager, how the team runs), or what the work does to you (stress, state of mind). Users cannot tell which room they are in.
Emmy: "พี่ไม่แน่ใจว่า คือที่เราทําอันเนี้ย คือเพื่อที่จะแบบว่า ให้เขาระบายออกมาเป็นเพื่อนคุย หรือว่าจะแบบปรับเปลี่ยนโครงสร้างขององค์กรเลยจริงจริง"
Tong, independently: "คิดว่าเป็นเพื่อนคุยคับ แต่ตอนนี้เพื่อนคุยดูอยากจะไปปรับโครงสร้างองค์กรเค้า"
Emmy's request if the answer is "friend": "อาจจะแบบเพิ่มการให้กำลังใจได้นะ ตอนตอบกลับ จะรู้สึกเป็นเพื่อนมากกว่า". The reply tone has to acknowledge before it analyses.
Current position: the engine serves both, and what we sell is what is inside a person's head at work. That settles the architecture. It does not settle the entry screen, and it is the entry screen the user meets.
Fix
Accept when
| problem | who | fix | accept when |
|---|---|---|---|
| Will not open in Chrome, had to use Safari | Neung | Find and fix the Chrome failure | Full flow completes in Chrome on a phone from a cold load |
| Wanted to carry on talking at the end but there was no microphone | Neung | There are two summary pages. The first has cards to tick and is still editable, so the mic belongs there. The last one has nothing to edit. Tong's call: leave the mic on throughout rather than reasoning about it per screen | Mic control present and working on every screen that still accepts a change |
| Summary ran about 11 pages and still left out problems he raised. It felt like it ignored some of what he said | Neung | Summarising in one pass makes the model drop everything except what it latched onto first. Run the analysis as a loop, build the summary from the loop, then let the person check it | Every distinct problem raised in a scripted walk appears in the summary |
| Speech recognition needs to be better, in Thai and in English | Neung | Improve transcription quality for both languages, not Thai alone. People code-switch mid-sentence | A scripted mixed Thai and English passage transcribes both without dropping the English |
| Text in the box flickered between old and new while speaking | Than | Freeze displayed text during active speech | Displayed text does not change between speech starting and the turn being captured |
| No intro explaining what this is or how to use it | Than | Add an onboarding screen before the first question | A screen precedes the first question board and explains the interaction |
| Swap-question prompt appeared mid-answer and confused him | Than | Suppress it while answering, or explain what it does | No swap prompt renders between a question being picked and the answer captured |
| Typo in the last box | Emmy | Proofread all UI strings | A native Thai reader signs off the string table |
| Wanted the names of people he mentioned used in questions | Neung | Reference named people from the transcript in generated questions | In a walk naming a person, a later question refers to them by that name |
Two testers praised the summarising and one said the opposite. That is not a difference of opinion, it is a limit of the product showing up.
| who | verdict | their words |
|---|---|---|
| Emmy one main thread | praised it | "เอไอตัวจับดีเลย" · "สรุปเรื่องโอเคมากกก" · "จับใจความเก่งอะ" · on the AI catching what she said and offering the next thread: "อันนี้ชอบมาก คือเลิศ" |
| Than answered shallowly on purpose |
praised the idea | "มีสรุปให้ รู้สึกดีนะ แบบสำหรับคนมีอะไรในหัวเยอะ ช่วยสรุป figure out ปัญหาให้" He gave it little to work with, so this is praise for the concept, not a test of it |
| Neung most problems, longest session, 30 turns |
said it failed | "ตอนนี้รู้สึกว่ามันยังจับปัญหาได้ไม่ค่อยดี" · summary came out around 11 pages and still "สรุปปัญหาออกมามันไม่ครบ" · "มัน ignore ปัญหาบางอย่างของเราไป" |
The pattern: it holds on one thread and breaks when there are several. Neung brought the most problems and got the worst comprehension. Even Emmy, who liked it, hit the edge of the same thing: "เอ๊ะ เหี้ย เยอะมาก ไม่รู้ต้องพูดอันไหนก่อน". This is why Neung asked for colour coding and a triage step, and it is the same root as problem 2.
Neung's two complaints have different causes, and the timing proves it
His session ran 7 minutes. The rate-limit refusals start at 5 minutes 36 seconds and run to the end. The first five and a half minutes were clean. So:
Settle it by fixing the quota and rerunning him. If the summary is still incomplete on a paid account, that is a real finding about the product.
Round one is not a baseline. One session was rate-limited, one tester deliberately withheld, and every session had a broken board. Round two is round one done properly, so do not report it as an improvement.
Gate: problems 2 and 5 must pass their acceptance before anyone is invited back. Six testers have already been spent once.
Pass mark: each tester says, without being prompted, that they knew when it ended and what the product offers.