AI 產品 · HCI 研究 · 網站製作

我做能幫助人開口、學習與連結的AI 產品

我是蔡秉叡 Ray,臺大資工人工智慧碩士班(HCI Lab)。從語音辨識研究走進人機互動,獨立開發並上架 AI 口說練習 App,也幫社群與課程做網站。

蔡秉叡 Ray 在 Google 大廳的彩色牆與 Google 字樣前比讚
NTU CSIE · HCI Lab
Mirror已上架 App Store

從語音研究走進人機互動,做出上架的產品

研究教我怎麼量測,產品教我怎麼取捨。下面是我目前的學歷與經歷。

看完整履歷英文・2 頁・PDF

研究經歷

2025/6 – 2026/3

中央研究院 資訊科學研究所 CKIP Lab

兼任研究助理

創創學程

19屆

第 19 屆學員(2026)

15人

矽谷探索學程首屆 15 人

臺大創意創業學程

第 19 屆學員(2026);也是臺大矽谷探索學程首屆 15 位學員之一。

2026 夏天

約7週

臺大矽谷探索學程學員

「國際創新創業探索與實踐:矽谷探索實務」:約 7 週在矽谷做市場研究與專家訪談,9/12 參加成果發表會。

看矽谷七週回顧 (在新分頁開啟)
夜晚的大合照
夜晚的大合照
辦公室大合照
辦公室大合照
優勝美地的日落
優勝美地的日落
在矽谷辦公室分享
在矽谷辦公室分享
在 Google 的團體自拍
在 Google 的團體自拍

影片 · YouTube · 10:36

我把矽谷七週剪成一支影片

Seven Weeks in Silicon Valley | Growing Into My Own Voice

按下播放後才會載入 YouTube。

在 YouTube 觀看 (在新分頁開啟)

把研究 與好奇心

做成 每天用得到的工具

上架的 App,與正在做的產品

Mirror 首頁:像素角色與五維聲紋條
Mirror 聲音指紋分析結果

已上架 App Store

Mirror – AI Speech Coach

AI 語音教練:錄下你的口語回答,Mirror 會轉錄、評分並重寫成更清楚、更有結構的版本,再用你自己克隆出來的聲音播放出來,讓「更好的自己」聽起來不遙遠。

  • 聲音簽章:30 秒建立個人語音模型
  • 模擬面試:英/中/日三語
  • 原始回答與優化版本並排回放
  • 聲紋雷達:五個維度的說話風格分析

獨立開發|iPhone|免費下載

在 App Store 查看(在新分頁開啟)

其他專案

點開每個專案,看它的畫面、狀態與連結。

Echo iOS · Chrome
Echo 產品介紹頁

會記住你「倒退 15 秒」的 Podcast 播放器:把你重聽的那一句留下來,判斷沒聽懂的原因(單字、連音、語速…),隔天變成你的口說練習。Chrome 版把 YouTube 字幕互動變成同樣的練習素材。

  • iOS:TestFlight 受邀測試中
  • Chrome 擴充功能:開發中
Rayality Projection Mapping Web · 開源
Rayality 編輯器的四角校正畫面與投影輸出

在瀏覽器裡做投影映射:匯入圖片或影片、拖動四個角校正透視,再從另一個視窗輸出到投影機。核心功能不需要任何金鑰;也可以自備 Google Veo 金鑰生成動態素材。

  • 線上可用・開源
MedBuddy Web · LINE bot
MedBuddy 照顧者網頁儀表板(示範資料)

給多重用藥長者與家人的用藥理解與照護交接原型:用規則式判讀整理藥物資訊,長者和照顧者透過兩支 LINE 共用一份照護紀錄;藥袋照片的辨識結果要經人工確認才會寫入。在一場 Build Challenge 中約 48 小時完成。

  • 原型・示範資料
Echo 產品介紹頁Rayality 編輯器的四角校正畫面與投影輸出MedBuddy 照顧者網頁儀表板(示範資料)

我做的網站

我的 AI 協作法,開源成一本手冊

開源專案

Orca 多代理協作手冊Orca Coordination Handbook

  • MIT 授權
  • 繁中+English
  • v1.0.0

我平常同時用好幾個 AI agent 做事:兩位總控分工督導、每份產物只有一位 writer、交付要附證據才驗收。我把這套做法寫成教學手冊與可直接複製的模板,繁中與英文完整對照,可下載 HTML、PDF 與 ZIP。

12 章+能力對照附錄・每種語言 13 份模板與 6 張圖GitHub 公開 API・2026-10-06

開源給大家的部分

  • MIT 授權:可以直接 clone,或在 GitHub 按「Use this template」建立自己的版本
  • 複製模板不需要 AI 帳號、API key 或安裝 Orca;初始化工具只複製檔案
  • 清楚區分 Orca 原生功能、自製參考工具、人工規約與未知事項

獨立的社群教學與模板,不是 Orca 官方產品,也不代表 Orca 背書;我私人使用的助理執行環境(資料庫、排程等)沒有開源。

GitHub · rick-ray-wldd

我的程式碼都在 GitHub

Echo、Rayality、MedBuddy 與這本手冊的原始碼都公開在這裡。

9GitHub 公開專案GitHub 公開 API・2026-10-06

看我的 GitHub (在新分頁開啟)

數據

兩站累計瀏覽次數
暫無資料(統計起日:暫無資料・自建計數服務(Cloudflare))
GitHub 公開專案
9(GitHub 公開 API・2026-10-06)
開源模板(每種語言)
13 份 × 2 語言(GitHub 公開 API・2026-10-06)
手冊章節
12 章+附錄(GitHub 公開 API・2026-10-06)
已上架 App Store 的 App
1 款(Mirror)(App Store 頁面・2026-10-06)
手冊圖解(每種語言)
6 張 SVG(GitHub 公開 API・2026-10-06)
手冊最新版本
v1.0.0(GitHub 公開 API・2026-10-06)
本站收錄作品
9 件(本站內容(App 與產品 4+網站 4+開源 1))

製作案例

這個網站,是用一套可重複的流程做出來的

我設計了一套讓 AI 程式代理照著走的網站製作流程:我提出需求、挑參考、做取捨並驗收;代理負責拆解參考、寫程式、跑檢查、錄影和產生收據。每一步都要有證據才往下走。

  1. 01 我決定

    需求+參考

    一句話寫下要什麼,附上 Pinterest 參考。代理先下載參考影片、逐格拆成效果規格。

    參考影格表、效果對照表

  2. 02 代理執行

    Brief

    把原話、成功條件、禁止事項(不捏造、不付費、不對外發送)和驗收方式寫成一份 brief,我確認後才開工。

    brief 文件

  3. 03 代理執行

    Skills+唯一 writer

    代理讀可重用的技能檔(流程、檢查清單、發布規則)。每份產物只指派一位 writer,另一位代理只負責審查。

    接單紀錄

  4. 04 代理執行

    設計與開發

    所有文案和事實集中在一個設定檔:同一份內容產生中英兩版,也產生作品集與這個履歷版。動效都有減少動態與無 JS 的靜態版本。

    原始碼 commit

  5. 05 自動檢查

    電腦/手機 QA

    headless Chrome 自動檢查 360 到 1440 六種寬度:溢出、console 錯誤、連結實際打開、對比、鍵盤、觸控、離開畫面就暫停。每次修改都重跑整套。

    QA 報告

  6. 06 自動檢查

    錄影+參考對比

    錄下電腦版和手機版的真實捲動,再把成品和 Pinterest 參考並排。就是下面這三支影片。

    三支影片

  7. 07 我決定

    審核與發布

    另一位代理審查同一個凍結版本;build 產生逐檔 SHA-256 清單。發布後匿名讀回線上檔案,雜湊一致才算完成。

    manifest、線上讀回

  8. 08 代理執行

    Discord 交付

    把三支影片和摘要送到團隊頻道,再讀回訊息與附件大小確認送達;送之前先查有沒有送過,避免重複。

    送達收據

實際用到的工具:Vite+TypeScript、GSAP、headless Chrome(puppeteer-core)、ffmpeg、Cloudflare Pages 與 Workers、多代理工作區、Discord bot。

成品影片

電腦版(1440×900 實際錄影)headless Chrome 真實時間錄製這個履歷版。
手機版(390px 模擬)390px viewport+iPhone UA 模擬,不是真機錄影。
和 Pinterest 參考的對比左:Pinterest 參考影片;右:本站數據環(390px 模擬)。

參考來源:Pinterest「Infograph Spinning Animation」 (在新分頁開啟)

  • 學的:斜放的橢圓環、前排放大後排縮小的透視、有厚度的白色圓角卡片、被環擋住一部分的大字。
  • 改的:配色與字體換成本站的;每個數字都是有來源的真實數據,不畫沒有資料的曲線;大字保持可讀;加上拖曳、點卡片看來源與靜態版。

影片按下播放才會下載。

兩段主要 prompt(範本,可直接複製)

第一段從需求做到可驗收的成品;第二段負責修改、發布與交付。這是整理後的範本,你可以填入自己的需求直接使用。

Prompt 1:從需求到可驗收的網站

輸入
一句需求、參考連結、可用的素材與事實來源
產出
brief、雙語網站、QA 報告、電腦/手機/對比影片、交付收據
驗收
六種寬度無溢出、console 0 錯誤、連結實際可開、每個事實都有來源
你是這個網站的唯一 writer。
需求:<一段話,例如「個人履歷網站,中英雙語,給招募方看」>
參考:<Pinterest 或網站連結>——先下載並逐格拆解,列出要學的效果和不學的部分。
素材與事實:只能用 <資料夾或文件>;查不到就留空,不要編。

請完成:
1. 寫 brief:成功條件、禁止事項、驗收方式。
2. 實作:文案和事實集中在一個設定檔;中英兩版;動效要有減少動態與無 JS 的靜態版本。
3. QA:360/375/390/414/820/1440 都沒有溢出;console 0 錯誤;連結實際打開;鍵盤和觸控可用。
4. 錄影:電腦 1440、手機 390(標明是模擬)、與參考並排對比。
5. 交付收據:改了什麼、每個事實的來源、QA 結果、還沒完成的事。

限制:不部署、不付費、不對外發訊息。

Prompt 2:修改、發布、交付

輸入
回饋清單、同一位 writer、發布授權
產出
修改後的凍結版本、SHA-256 清單、線上讀回、送達收據
驗收
審查者對同一版本通過;線上檔案雜湊一致;影片送達且可讀回
依下面的回饋修改同一個網站(你仍是唯一 writer):
<回饋清單>

1. 每一項改完都附截圖或測試證據;重跑完整 QA 與錄影。
2. 凍結:commit、乾淨 build、逐檔 SHA-256 清單(LC_ALL=C 排序)。
3. 交給審查者對同一個版本互審;通過後由唯一的部署者發布。
4. 發布後匿名讀回:中英首頁、資產雜湊、關鍵文字。
5. 把電腦、手機、參考對比三支影片送到 <頻道>,讀回附件名稱與大小留存收據;送之前先確認沒有送過。

回報:commit、manifest、測試數、影片路徑、還沒驗證的事。

可攜的流程指南(Skills)

Skills(技能檔)是寫給 AI 代理讀的工作說明書:流程、閘門、檢查清單和界線都寫清楚,每次做網站就走同一套步驟。這份是我另外寫的公開版,可以直接放進你自己的代理工作區,也可以自由改寫。

多代理分工與驗收的完整做法,見我的開源協作手冊。 GitHub (在新分頁開啟)

實際時間紀錄

  1. 需求 brief 與參考素材落檔
  2. 初版:雙語靜態網站+自動 QA+交付收據
  3. 6 輪修改(照片、影片區、首屏照片、作品排列、聯絡表單等)
  4. 正式網站發布並記錄線上驗證
  5. 數據環、開源手冊卡、兩站共用瀏覽計數(含 1 輪修正)
  6. 之後:這個履歷版與製作案例

時間取自 git commit 紀錄(2026-10-06,台北時間)。這些是時間點,不是工作時數;中間包含等待與休息,初版和之後的修改分開列。

聯絡我

留個話,我會回覆你

寫下怎麼稱呼你、怎麼聯絡你,以及想聊的事就好。

例如:王老闆、陳小姐

怎麼聯絡你(至少填一項)
想聊什麼(可以多選)

例如:想聊聊 AI 產品相關的職缺

留言只存在網站後台、僅 Ray 本人可查看,用於回覆你。

蔡秉叡 Ray Tsai

Research & Product Focus

Multimodal human-computer interaction, sports coaching technology and voice AI. Combines basketball tactical knowledge, study design and product development to explore how personalized audio can support learning and real-world practice.

Education

National Taiwan University (NTU)Feb 2026 – Present

M.S. student, Computer Science and Information Engineering, Artificial Intelligence Program

  • HCI Lab, advised by Prof. Mike Y. Chen(在新分頁開啟); multimodal interaction, emotional voice and AR/VR interfaces.
  • Spring 2026 GPA: 4.04/4.3 (15 credits). A+ in Multimodal HCI, Natural Language Processing, Artificial Intelligence, and Web Programming; additional coursework in Big Data Systems.

National Yang Ming Chiao Tung University (NYCU)Sep 2021 – Dec 2025

B.S. in Electrophysics

  • GPA: 4.1/4.3 over the final four semesters. Graduate-level coursework: Machine Learning for Signal Processing (A+), Speech Signal Processing (A+), Deep Learning (A), and AI and Law (A+).

Research Manuscript Under Review

CoachCast - Player-Specific Audio for Basketball Practice2026

Coauthor | Submitted to ACM CHI 2027; under review

  • Contributed basketball tactical guidance and recommendations, experimental design, and data annotation to a study of concurrent, player-specific audio guidance during coach-led team practice.
  • The project examines how personalized audio supports coordinated player actions alongside team-wide coaching, including trade-offs around guidance timing and player autonomy.

International Entrepreneurship Experience

NTU Silicon Valley Exploration ProgramJul – Aug 2026

Selected participant, inaugural 15-student cohort | San Francisco Bay Area, USA

  • Completed seven weeks of startup exploration, industry visits and founder exchanges; presented product progress and reflections through weekly NTU Day sessions at Startup Island TAIWAN.
  • Brought hands-on product feedback to IrisGo's cofounder and CTO: reproduced a Windows integration detection issue and narrowed likely causes to environment, executable detection and launch behavior.
  • Discussed creator voice-generation workflows with Influenxio's founder; translated ecosystem feedback into Echo's language-learning direction and a subsequent expressive-voice editing prototype.
  • Presented at Pre-Demo Day and developed MedBuddy for AI Fund's Engineer in Residence Build Challenge; continued product development and professional relationships after returning to Taiwan.

Selected Products

Mirror - AI Speech Coach(在新分頁開啟)Apr 2026 – Present

Solo Founder | App Store release; 2026 NTU AI Builders Challenge entry

  • Built an end-to-end coaching app that analyzes spoken answers and replays revised responses in the user's cloned voice, supporting Traditional Chinese, English and Japanese.
  • Implemented voice calibration, streaming TTS, mock interviews and onboarding with Expo/React Native, Next.js and Supabase.
  • Reached 718 organic signups; shipped voice calibration, practice and feedback as a complete mobile learning workflow.

Echo - Contextual Language LearningJul 2026 – Present

Cofounder & Product Developer

  • Built a podcast-learning prototype linking playback and transcripts with vocabulary capture and replay; uses rewind behavior as a signal for comprehension support and follow-up shadowing practice.
  • Lead product development with a cofounder (NTU Information Management), who handles native Swift, Dynamic Island, system playback controls and practice cards.
  • Exploring browser-extension workflows for learning from video and other media beyond podcasts; developing Echo through NTU entrepreneurship coursework.

MedBuddy - Medication-Care Prototype(在新分頁開啟)Jul 2026

AI Fund Engineer in Residence Build Challenge submission

  • Delivered a web dashboard and LINE-based care prototype during a 48-hour build challenge, connecting older adults and family caregivers; collaborated on LINE transport and reminder scheduling.
  • Separated deterministic medication-rule verdicts from AI explanations and documented rule provenance. Project validation recorded 316 passing tests; prototype has not undergone clinical validation.

Emotional Voice - Creator Voice EditingSep 2026 – Present

Independent Developer | Single-user local prototype

  • Built a take-comparison workspace with manual and AI-assisted editing, change previews, undo, candidate listening and WAV export; 11 integration tests and desktop/mobile checks passed.
  • Compared English, Chinese and mixed-language voice-clone recordings through controlled generation and exploratory listening. Investigating creator advertising workflows informed by discussions with Influenxio.

Research Experience

Research Assistant (Part-time), Academia SinicaJun 2025 – Mar 2026

CKIP Lab, Institute of Information Science | Dr. Wei-Yun Ma

  • Worked on multimodal language models and emotion-aware speech generation; contributed progressive SpecAugment scheduling and dialect-aware conditioning to the 2025 Formosa Speech Workshop Challenge team's low-resource Hakka ASR work.
  • Surveyed Kimi-Audio for joint emotion recognition and controllable speech generation.

Research Assistant, National Tsing Hua UniversityMar 2023 – Jul 2025

Institute of Learning Sciences & Technologies

  • Co-developed a multimodal AI picture-book platform combining speech synthesis and image generation for language learning and cognitive therapy; taught 12 hours per semester of generative AI and n8n/LangChain workshops.

Shadow Your Perfect SelfMay – Jun 2026

NTU Multimodal HCI course research project

  • Built and evaluated a three-condition self-voice shadowing platform with 12 participants; observed preference for self-voice models, while objective measures did not establish pronunciation gains.

Athletics, Leadership & Service

  • NTU Varsity Basketball Team, Forward (Fall 2026 - Present): joined this semester; registered for the 2027 北市大盃 tournament.
  • C-Level Basketball Coaching License (Taiwan): certified basketball coach.
  • NTU Creativity and Entrepreneurship Program, 19th cohort (2026); selected for High-Tech Entrepreneurship and Operations, a 12-student course (Fall 2026).
  • Public Relations Officer, NYCU Electrophysics Student Association (2022-2023): hosted 9 career seminars with 700+ total attendees and speakers from semiconductor and photonics companies.
  • Awards: Runner-Up, Law Broker GPT PLUS+ AI Competition (Jan 2024); Honorable Mention, 美術竹梅 Hackathon (2023).
  • ITRI Research Intern (Oct 2022 - Jun 2023): Python tools for automated data collection. DuPont Taiwan IT Intern: corporate IT workshops and network maintenance.

Ping-Juei (Ray) Tsai | September 2026

公開版本,未列電話與 email;想聯絡我,請用本站的聯絡表單。 前往聯絡表單