AI products · HCI research · Web

I build AI products that help people speak, learn and connect.

I'm Ray Tsai, in the AI master's program at NTU CSIE (HCI Lab). I started in speech-recognition research, shipped an AI speaking-practice app on my own, and build websites for communities and programs.

Ray Tsai giving a thumbs-up in front of the Google sign and colorful lobby wall
NTU CSIE · HCI Lab
MirrorLive on the App Store

From speech research to shipped products

Research taught me how to measure; products taught me what to cut. Here is where I study and what I have done.

Education

National Taiwan University

Dept. of Computer Science & Information Engineering — AI master's program, HCI Lab

Research

2025/6 – 2026/3

CKIP Lab, Institute of Information Science, Academia Sinica

Part-time research assistant

Speech recognition challenge

2nd

Track 1 — 2nd place

3rd

Track 2 — 3rd place

Formosa Speech Recognition Challenge 2025

Team entry (student division); the team's paper was published at ROCLING 2025.

Summer 2026

≈7weeks

NTU Silicon Valley exploration program

About seven weeks in Silicon Valley doing market research and expert interviews, followed by the Demo Day on Sep 12.

Read the seven-week recap (opens in a new tab)
Group photo at night
Group photo at night
Group photo at the office
Group photo at the office
Sunset at Yosemite
Sunset at Yosemite
Sharing at an office in Silicon Valley
Sharing at an office in Silicon Valley
Group selfie at Google
Group selfie at Google

Film · YouTube · 10:36

I turned my seven weeks into a film

Seven Weeks in Silicon Valley | Growing Into My Own Voice

YouTube loads only after you press play.

Watch on YouTube (opens in a new tab)

Turning research and curiosity

into tools people use daily

A shipped app, and products in progress

Mirror home screen with pixel character
Mirror voice fingerprint result

Live on the App Store

Mirror – AI Speech Coach

An AI speech coach: record your spoken answer, and Mirror transcribes, scores and rewrites it for clarity and structure — then plays the better version back in your own cloned voice.

  • Voice signature in 30 seconds
  • Mock interviews in English, Chinese, Japanese
  • Side-by-side replay of original vs. improved
  • Five-axis voice radar

Built solo · iPhone · Free

View on the App Store(opens in a new tab)

More projects

Open a project to see its screens, status and links.

Echo iOS · Chrome
Echo product page and a lock-screen practice prompt

A podcast player that remembers every 15-second rewind: it keeps the sentence you replayed, diagnoses why you missed it (vocabulary, linking, speed…), and turns it into tomorrow's speaking practice. The Chrome version does the same with YouTube captions.

  • iOS: invite-only TestFlight beta
  • Chrome extension: in development
Rayality Projection Mapping Web · Open source
Rayality editor corner-pinning and projector output

Projection mapping in the browser: import images or video, drag four corners to correct perspective, and send the output to a projector from a separate window. No API key needed for the core workflow; optional Google Veo generation with your own key.

  • Live · open source
MedBuddy Web · LINE bot
MedBuddy caregiver dashboard (demo data)

A prototype for medication understanding and care handoffs for older adults on many medicines: rule-based checks, and two LINE accounts sharing one care record between elder and caregiver. Medication-bag photo readings must be confirmed by a person before they are saved. Built in about 48 hours during a build challenge.

  • Prototype · demo data
Echo product page and a lock-screen practice promptRayality editor corner-pinning and projector outputMedBuddy caregiver dashboard (demo data)

Websites I've built

My multi-agent workflow, as an open handbook

Open source

Orca Coordination Handbook

  • MIT license
  • English + 繁中
  • v1.0.0

I run several AI agents at once: two coordinators split supervision, every deliverable has exactly one writer, and nothing is accepted without evidence. I wrote that workflow up as a handbook with copy-ready templates, complete in English and Traditional Chinese, with HTML, PDF and ZIP downloads.

12 chapters + capability appendix · 13 templates and 6 diagrams per languageGitHub public API・2026-10-06

What's open

  • MIT licensed: clone it, or press “Use this template” on GitHub to start your own copy
  • Copying the templates needs no AI account, API key or Orca install; the helper only copies files
  • Clearly separates Orca-native features, private reference tooling, written conventions and unknowns

An independent community handbook and template set — not an official Orca product or endorsement. My private assistant runtime (database, scheduler, etc.) is not open-sourced.

GitHub · rick-ray-wldd

My code lives on GitHub

Source for Echo, Rayality, MedBuddy and this handbook is public there.

9Public GitHub reposGitHub public API・2026-10-06

Visit my GitHub (opens in a new tab)

By the numbers

Numbers that spin — and can be checked

Page views are one counter shared by both of my sites; the rest come from public GitHub data and this site's own content. The list below shows current values and sources.

Current values and sources

Combined page views
No data yet
Counting since: No data yet・Self-hosted counter (Cloudflare)
One counter shared by my portfolio and résumé sites. Reloading or switching language in the same tab counts once. Not unique visitors.
Public GitHub repos
9
GitHub public API・2026-10-06
Open templates (per language)
13 × 2 languages
GitHub public API・2026-10-06
Handbook chapters
12 + appendix
GitHub public API・2026-10-06
Apps live on the App Store
1 (Mirror)
App Store listing・2026-10-06
Handbook diagrams (per language)
6 SVG each
GitHub public API・2026-10-06
Latest handbook release
v1.0.0
GitHub public API・2026-10-06
Handbook GitHub stars
0
GitHub public API・2026-10-06
Projects on this site
9
This site (4 apps & products + 4 websites + 1 open source)
Handbook license
MIT
GitHub public API・2026-10-06

Case study

This site was built with a repeatable workflow

I designed a website workflow that AI coding agents follow step by step: I set the goal, pick references, make the calls and sign off; the agents break down references, write the code, run the checks, record the videos and write receipts. No step moves on without evidence.

  1. 01 I decide

    Need + reference

    One sentence on what I want, plus a Pinterest reference. The agent downloads it and breaks it into an effect spec, frame by frame.

    Frame sheet, effect map

  2. 02 Agent runs

    Brief

    My words, success criteria, hard limits (no made-up facts, no spending, no outbound messages) and acceptance tests go into one brief that I approve.

    Brief document

  3. 03 Agent runs

    Skills + one writer

    Agents load reusable skill files (process, checklists, release rules). Each deliverable gets exactly one writer; a second agent only reviews.

    Assignment record

  4. 04 Agent runs

    Design & build

    All copy and facts live in one content file, which generates both languages — and both the portfolio and this résumé. Every animation has a reduced-motion and no-JS static fallback.

    Source commits

  5. 05 Automated

    Desktop / mobile QA

    Headless Chrome checks six widths from 360 to 1440: overflow, console errors, live links, contrast, keyboard, touch, pausing off-screen. The full suite reruns on every change.

    QA report

  6. 06 Automated

    Recordings + comparison

    Real-time scroll recordings on desktop and mobile, plus the result side by side with the Pinterest reference — the three videos below.

    Three videos

  7. 07 I decide

    Review & publish

    A second agent reviews the same frozen version; the build gets a per-file SHA-256 manifest. After publishing, live files are read back anonymously and must match.

    Manifest, live read-back

  8. 08 Agent runs

    Discord delivery

    The three videos and a summary go to a team channel; the message and attachment sizes are read back as a receipt, after checking it wasn't already sent.

    Delivery receipt

Tools actually used: Vite + TypeScript, GSAP, headless Chrome (puppeteer-core), ffmpeg, Cloudflare Pages and Workers, a multi-agent workspace and a Discord bot.

Recordings

Desktop (1440×900 real-time recording)Real-time headless Chrome recording of this résumé site.
Mobile (390px simulation)390px viewport with an iPhone user agent — a simulation, not a real device.
Side by side with the Pinterest referenceLeft: the Pinterest reference; right: this site's data ring (390px simulation).

Reference: “Infograph Spinning Animation” on Pinterest (opens in a new tab)

  • Kept: the tilted elliptical ring, near-large/far-small perspective, thick rounded white cards, and a big wordmark partly hidden by the ring.
  • Changed: this site's colors and type; every number is real and sourced, with no data-free curves; the top and bottom words stay fully readable; added pause, next, drag and a static version.

Videos download only when you press play.

Two main prompts (templates you can copy)

The first goes from a need to a reviewable build; the second handles revisions, publishing and delivery. These are cleaned-up templates — fill in your own details.

Prompt 1: from a need to a reviewable site

Input
A one-line need, a reference link, the allowed facts and assets
Output
Brief, bilingual site, QA report, desktop/mobile/comparison videos, receipt
Acceptance
No overflow at six widths, zero console errors, live links, every fact sourced
You are the only writer for this website.
Need: <one paragraph, e.g. "a bilingual résumé site for recruiters">
Reference: <Pinterest or website link> — download it and break it down frame by frame; list what to borrow and what not to.
Facts and assets: use only <folder or document>; if something can't be verified, leave it out.

Deliver:
1. A brief: success criteria, hard limits, acceptance tests.
2. The build: all copy and facts in one content file; two languages; every animation with reduced-motion and no-JS static fallbacks.
3. QA: no overflow at 360/375/390/414/820/1440; zero console errors; links actually open; keyboard and touch work.
4. Recordings: desktop 1440, mobile 390 (labelled as a simulation), and a side-by-side with the reference.
5. A receipt: what changed, the source of every fact, QA results, open items.

Limits: no deploying, no spending, no outbound messages.

Prompt 2: revise, publish, deliver

Input
A feedback list, the same writer, permission to publish
Output
A frozen revision, SHA-256 manifest, live read-back, delivery receipt
Acceptance
Reviewer passes the same version; live hashes match; videos delivered and read back
Revise the same website using the feedback below (you are still the only writer):
<feedback list>

1. Attach a screenshot or test evidence for every item; rerun the full QA and recordings.
2. Freeze: commit, clean build, per-file SHA-256 manifest (sorted with LC_ALL=C).
3. Hand the same version to a reviewer; once it passes, the single deployer publishes it.
4. After publishing, read back anonymously: both language home pages, asset hashes, key text.
5. Send the desktop, mobile and comparison videos to <channel>, read back attachment names and sizes as a receipt; check first that they weren't already sent.

Report: commits, manifest, test counts, video paths, anything not yet verified.

Portable workflow guide (skills)

Skills are instruction files written for AI agents: the process, gates, checklists and limits, so every site follows the same steps. This is a public version I wrote separately — drop it into your own agent workspace and adapt it freely.

For the full multi-agent coordination method, see my open-source handbook. GitHub (opens in a new tab)

What actually happened, and when

  1. Brief and reference material committed
  2. First build: bilingual static site + automated QA + receipt
  3. 6 revision rounds (photos, film section, hero photo, works order, contact form, etc.)
  4. Production site published and verified live
  5. Data ring, open-source handbook card, shared view counter (incl. 1 fix round)
  6. Then: this résumé version and case study

Times come from the git commit history (2026-10-06, Taipei time). They are timestamps, not hours worked — they include waiting and breaks. The first build and later revisions are listed separately.

Contact

Leave a note, I'll get back to you

Just tell me your name, how to reach you, and what you'd like to talk about.

e.g. Mr. Wang

How can I reach you? (at least one)
What's it about? (choose any)

e.g. I'd like to talk about an AI product role

Your details are only used to reply to you.

Or email me directly at allcare.rickray@gmail.com