Teacher, the reason you're struggling with voice input is the 'microphone'—The last mile of AI voice input, starting with the ¥2,780 EarPods
Whether it's Aqua Voice, Gemini, or the iPhone's default voice input—if you change the microphone, everything becomes a 'different experience'
[Understand in 10 seconds] The conclusion of this article
✅ The real reason you struggle with voice input is the 'microphone.' Just by bringing the ¥2,780 EarPods (USB-C) close to your mouth, the recognition accuracy becomes completely different from the built-in microphone.
✅ The top choice for AI is Aqua Voice (approx. $8/month when billed annually, as of June 2026). However, for teachers who cannot install it on school computers, the school-approved Gemini microphone button is a sufficient alternative.
✅ Even if Bluetooth is prohibited on school computers, wired USB-C EarPods are likely to pass through via USB Audio Class (please confirm with your IT department for final verification).
✅ The ultimate setup for serious expansion is Shokz OpenComm2 UC (¥27,880) at home × Aqua Voice. Including the time between chores, your teaching paperwork can be handled almost entirely by 'just talking.'
I am writing this for teachers like you
Teachers who have started using Aqua Voice or Superwhisper but feel that 'it doesn't recognize as well as I thought' while still using the PC's built-in microphone.—This article is aimed at you for more than half of its content. Changing the microphone will change your world.
Mainly junior high school teachers. Teachers who do not have the authority to install new software on school computers and feel that voice input is impossible for them.
Even elementary school teachers who feel that 'typing can't keep up' in the staff room.
Teachers who have heard the names Aqua Voice or Superwhisper but have stopped due to the barriers of payment or setup.
Teachers who have Google for Education at school and have used Gemini.
As a related article,'The Ultimate AI Environment for Teachers: Cursor × Claude Code × Voice Input' covers the overall picture of an AI environment including voice input. You don't need to have read it. This article focuses on the step before that, the topic of 'What you talk to the AI with = the microphone.' It doesn't matter if the software is Aqua Voice or Gemini.
Initial note: I will write about the handling of personal information only here.
The 'voice input' discussed in this article also touches on tasks involving personal information such as students' names, grades, and observations. As a premise,do not speak out loud personal names or information that can identify individuals. For tasks requiring proper nouns, speak only using numbers or context during dictation, and replace them with the proper nouns manually after the AI formats the text. This is the same practice used in medical settings when discussing cases as 'Mr. A'.
I will not repeat this warning in the rest of the text. The axis of judgment will be based on what I have declared here just once.
The core of voice input is the 'microphone'—Buy ¥2,780 EarPods before you get lost in choosing software.
The order is reversed.
The first thing a teacher who thinks 'I want to improve the accuracy of voice input' does is read articles about Aqua Voice or Superwhisper, stop because they can't bring themselves to pay $8 a month, or feel 'it doesn't recognize as well as I thought' even after paying. This is the biggest waste.
The reason is simple, and in my experience,voice input accuracy is close to '80% microphone, 20% software.' No matter how much the AI side evolves, if the input audio is 30–50cm away from the laptop's built-in microphone, misconversions will not decrease. Even if you pay $8 a month for Aqua Voice, if you are still using the built-in microphone, you are not even utilizing half of its performance.
The leading software on the AI side is certainly Aqua Voice. For those who can afford the $8 monthly subscription, it is currently the most powerful voice input software available. If you are going to do voice input seriously, this is the logical destination eventually. However, in reality, only a minority of teachers are able to install voice input software on their school-issued PCs.
School-issued PCs now typically use a whitelist system (a system where only pre-approved software can be installed) with administrative privileges stripped, so even if you double-click the Aqua Voice exe, the installation will not run.
Fortunately, even if you cannot use Aqua Voice on your school PC, Gemini, which is included in the Google for Education suite used by many schools, has a microphone button standard in the chat box. The AI officially approved by the school has a voice input button implemented from the start. It doesn't have the same accuracy as Aqua Voice, but the moment you bring the EarPods to your mouth, it crosses the line into 'sufficiently practical'.
In short, it's like this.
Ideal form (home, personal PC, personal smartphone): Aqua Voice + EarPods (or Shokz UC)
Realistic configuration for school PCs: Gemini microphone button + EarPods
The core common to both: Bringing the microphone to your mouth with the ¥2,780 EarPods (USB-C)
Before worrying about the type of software, first buy the ¥2,780 microphone. This is the correct order.
The '3 Misconceptions' that cause teachers to get stuck with voice input
I will break down the reasons why it doesn't work into three points. All of them are structural misunderstandings.
Misconception 1: 'You can't do decent voice input unless you install paid software like Aqua Voice.'
In terms of accuracy alone, it is true that Aqua Voice is the best. If you can afford to pay $8 a month, choosing this in the long run is the logical path. I think so too.
However, in most cases, Aqua Voice cannot be installed on school PCs. This is because whitelist management for school-issued PCs has become standardized, and this is not the individual teacher's responsibility, but a structural issue.
That is where the microphone icon in the Gemini chat input field becomes a lifeline. If you press this, voice input starts from that moment. Availability depends on conditions such as Workspace for Education's Education Plus or Teaching and Learning add-on, as well as administrator settings, so it is best to check the situation at your school with the IT person once ( click here for Google's official explanation).
The Gemini microphone button does not have the same accuracy as Aqua Voice. But, the moment you bring the EarPods to your mouth, it crosses the line into 'sufficiently practical'. Just by changing the position of the microphone, the difference from the built-in microphone is obvious. 'Since Aqua Voice is impossible on school PCs, I'll get by with Gemini + EarPods'—this is the first realistic step.
Misconception 2: 'Bluetooth is prohibited, so I can't bring in a microphone for voice input.'
As many teachers have realized, Bluetooth is disabled on school PCs. Some municipalities and school PCs also disable Bluetooth, so please check your school's device management policy once.
However, it is too early to give up here. Wired USB Audio Class devices are often managed separately from the USB storage prohibition policy. It is common for them to be designed to pass on many school PCs as a 'class that does not involve data writing.' EarPods USB-C fall into this class.
However, there are large differences between municipalities. Since there are municipalities that completely block USB devices, the safe approach is to try plugging your personal earphones into the school PC's USB port once before ordering to see if they are recognized as a 'USB Audio Device'. At the same time, confirm with the IT person by asking, 'Can I use a USB audio device (HID/Audio class)?' If the IT person cannot answer immediately, please use the application template in the appendix mentioned later.
Misconception 3: 'If a teacher is going to use voice input, they need a high-performance headset like Shokz from the start.'
This is also wrong.
For inputting in a low voice at the staff room level, EarPods are sufficient. The remote control unit has a built-in microphone, and when worn, it sits 10–15 cm below your mouth. If you cup your hand around your mouth, it will pick up even mumbled voices well enough.
With ¥2,780 wired earphones, low-voice operation in the staff room becomes viable. Bone conduction headsets like the Shokz OpenComm2 UC are better suited as a second device kept at home for full-scale use—the final configuration of Aqua Voice + Shokz— once you are used to it. If you start with a Shokz as your first device, you will hit a different wall: it is too big and too conspicuous for use in the staff room.
Bonus paradigm shift: Assume the AI side will handle misrecognitions
Aqua Voice claims high accuracy. In my experience, it has fewer misrecognitions than Gemini's voice input. However, even if there are some misrecognitions with the Gemini microphone button, Gemini itself will understand the context and format it. If you continue speaking after fillers like 'um' or 'uh' by saying 'fix this part like this,' the AI will reflect that and rewrite it.
The first step to increasing accuracy is simply to 'bring the microphone to your mouth.' You can worry about the final few percent after paying for Aqua Voice. If you change the microphone, even Gemini will perform differently than with the built-in microphone.
The minimum configuration to start today on your school PC—EarPods USB-C ¥2,780 × Gemini (The standard approach for those who cannot install Aqua Voice)
This is a realistic minimum configuration for teachers who cannot install Aqua Voice on their school PCs. Decide to use the ideal setup (Aqua Voice + EarPods) at home, and settle for Gemini microphone button + EarPods at work.
Step 0: Check if the Gemini microphone button is enabled for your school account (30 seconds)
The first line. If you skip this, you will drop out.
Open Gemini with your school Google account and check if a microphone icon is displayed at the bottom right of the chat input field. If it is displayed, a 'Do you want to allow access to the microphone?' dialog from Chrome will appear when you click it for the first time, so select 'Allow'.
I will provide alternative paths in case it is not displayed.
Open Gemini with your personal Google account and try the feel of the microphone button once (return to your school account for actual operation)
Install the Gemini app on your iPhone and get a feel for the voice button on your personal smartphone first
In parallel, ask the teacher in charge of IT once, 'Could you please enable voice input for Gemini on the teacher account?' (You can work with the two methods above while waiting for a reply)
Step 1: Order EarPods (USB-C) (2 minutes)
Order EarPods (USB-C) for ¥2,780 from the official Apple Store. The official Apple store delivers as early as the next day.
Type: In-ear
Connection: Wired USB-C connector (no driver required, recognized as a USB Audio Device)
Remote: Volume, playback, calls, and built-in microphone
Compatibility: Windows/Mac/iPad/iPhone 15 or later with USB-C
If your school PC only has a USB-A port, buy a USB-C to USB-A adapter as well. Search for 'USB-C female USB-A male adapter audio compatible'. You can find them from Anker or ELECOM for a few hundred to around ¥1,500 (model numbers change frequently, so choose the latest version).
Lightning users with an iPhone 14 or earlier can use stock Lightning EarPods or a Lightning-to-3.5mm adapter plus wired earphones as a substitute.
Step 2: How to input in a low voice in the staff room
Put on the EarPods.
Hold the remote control part in your left hand
Lift it to a position 10–15 cm below your mouth
Cup your palm over the remote control like a dome (no gaps)
Speak in a whisper close to your breath
Due to the proximity effect, your voice will be picked up loudly by the microphone, while surrounding noise will be relatively quiet. Your colleague sitting next to you will barely hear it.
For the first time, try it early in the morning around 7 AM or after 8 PM after school, when the staff room is almost empty, to reduce the psychological hurdle to zero. Once you get used to it, blend it into the noise before the morning assembly or during breaks. It also makes things easier for both of you if you tell the teacher next to you, 'I'm testing out voice input'.
Step 3: Have Gemini format it
Press the microphone button and speak your draft for the report. Do not say proper names; just speak the situation and your impressions.
'Second term report. A scene where they worked persistently on the geometry unit in math. At first, they stopped because they couldn't draw auxiliary lines, but after listening to a friend's explanation during group work, they were able to find their own way to solve it. Even during presentations, they were able to explain it in their own words with confidence. Uh, also, I saw them calling out to younger students during lunch cleanup.'
When you finish speaking, instead of pressing the stop button, continue by giving this instruction with your voice.
'Format the content I just said into a report style for a report card, about 40 characters per sentence, in two sentences. Remove the fillers.'
An example of the formatting Gemini will return:
'In the math geometry unit, although they struggled with auxiliary lines, they found their own solution through dialogue with friends and were able to present it with confidence. During lunch cleanup, they were seen proactively calling out to younger students, and their compassionate actions shone through.'
Proper names can be replaced later in bulk using Word's find and replace. Even if there are misrecognitions, if you just say, 'Change "auxiliary lines" to "logical thinking process"', Gemini will reflect that. Everything, including formatting instructions, is completed entirely by voice. Reducing the time you spend returning to typing by even one second—this is the core of the workflow I want to recommend in this article.
Once you get used to it, you can generate the material for one report in 2–3 minutes. Don't get discouraged if the first three take time, assuming it will take 3–5 times longer for the first attempt due to the technique, correcting misrecognitions, and replacing proper names.
It's not just for report card comments—everything from grade-level newsletters and teaching material preparation to organizing sticking points can be handled entirely by voice.
I'll write down my own actual workflow. Here is what is actually running with 'Voice Input × Gemini'.
Drafting grade-level newsletters: I dictate a rough draft of this week's events and next week's schedule during a 5-minute gap. By just continuing to speak to Gemini, saying 'Format this for parents, with each item being 2-3 lines,' it gets elevated to a layout ready for distribution.
Generating material for report card comments: I have created a GEM (a custom bot within Gemini) specifically for report card comments. Without including proper nouns or private information, I just speak about the situation, my impressions, and the direction of growth, and it returns a draft of the comment with a consistent writing style. All that's left is to manually fill in the proper nouns and make minor adjustments.
I wrote about how to create a report card comment GEM in 'Leaving teacher report card comments to Gemini's 'GEM'—How to create your own personal report card AI in 5 minutes'.
If you are new to the concept of report card AI, it is easier to start by reading 'Getting AI to give you a 60-point draft for report card comments' first.Organizing teaching sticking points: While looking at teaching materials, I just speak, 'Make this explanation a bit easier to understand' or 'Use an analogy for this,' and Gemini returns three or four alternative suggestions. Even in the 10-minute break between classes, this workflow works.
Full-scale lesson and material preparation at home: This is where the Shokz OpenComm2 UC comes in (described later). Basically, I just keep talking, and I only manually correct the accuracy of detailed historical or legal terms—that is the reality of my work at home.
Once 'Voice Input × AI Formatting' starts running, a significant portion of a teacher's text-based tasks, not just report card comments, will be replaced by this workflow.
The second battlefield of gap time—how to make the most of your commute, 10-minute breaks, and 20-minute breaks
For schools where no software installation is allowed on school computers, schools where the Gemini microphone button is disabled by administrator settings, or municipalities that block all USB devices—you don't have to give up.It is outside of the school computer that the time where voice input can truly show its potential lies.
The common equipment is the same. For iPhone 15 and later, use the USB-C EarPods directly (for iPhone 14 and earlier, use the Lightning version or a Lightning-to-3.5mm adapter). With the hand opposite the strap, or at your desk, lift the remote to your mouth, cover it, and speak in a low voice. The input destination can be the Gemini mobile app, the standard iOS voice input, or Aqua Voice installed on your personal smartphone.
Commuter train—as long as it's not packed, it's perfectly usable.
I personally did this on my commute. There is the circumstance that the train heading away from the city (opposite the direction of congestion) was relatively empty. Even so, speaking in a low voice inside the train while covering the remote with my hand at my mouth was sufficient for practical use.
You might be told that 'talking on the train is bad manners.' However, when you actually try it, the microphone picks up your voice perfectly even at a low volume where people around you absolutely cannot tell what you are saying. In my experience, it's fine as long as it's not a packed train. Even on days when I'm standing, I can at least dictate the structure of a memo into a note app while leaning against a handrail.
Of course, 'I want to read a book during my commute,' 'I want to look at the scenery,' or 'I want to zone out' are all perfectly fine. Please use it according to your own style. It is enough to have the stance that the commuter train is just one of the options available. Those who want to read can read, and those who want to use it can use it. That's fine.
10-minute and 20-minute breaks in the staff room—this might actually be the main event.
More convenient than the commuter train are the fragmented times within the school.
Between morning preparations, drafting a grade-level newsletter at my desk for 2 minutes
During a 10-minute break, having Gemini list 3 things that might be sticking points for the next class
During a 20-minute break, dictating a memo for afternoon parent communication in 1 minute
For 30 minutes after school, generating material for 5 students' report card comments
You can utilize those spare moments when you previously thought, 'I can't get started because my typing isn't fast enough,' by using voice input.Make full use of 10-minute or 20-minute breaks to keep going—I believe this is the real deal for voice input for teachers. Whether or not you use your commute on the train is a matter of preference, but spare time in the staff room is a shared resource available to everyone.
Continuing a draft created on the commuter train using the school office PC—The email transfer route
If you want to continue a draft created on your iPhone during your commute on your school office PC:Check once if 'emailing text data that does not contain personal information to your school address' is permitted under your operational guidelines. whether it is allowed.
My school uses this operation. I send the draft I made on my iPhone to my own school email address. When I arrive at school, I receive it on the office PC and edit the rest.
However, this is a point of contention that varies greatly by municipality. There are indeed large boards of education that restrict sending to school email addresses from personal devices under BYOD regulations. Please search for the clauses regarding 'sending to official addresses from personal devices' and 'handling of text data on personal devices' in your school's information security implementation procedures before taking action. Limit this to general documents that do not contain personal information (event notices, draft communications, teaching materials). Do not handle student names or grades on personal devices. This is a line you cannot cross.
For beginners, I recommend starting with Aqua Voice—For teachers who are too lost to move forward
When asked, 'Which voice input software should I choose in the end?', I answer, try starting with Aqua Voice first. There are many options: Superwhisper, Typeless, iOS standard voice input, and ChatGPT's voice. Once you get used to it, you can choose something else that suits you better. However, for teachers, the biggest waste is getting stuck while you're still undecided.
There are only three reasons why I recommend Aqua Voice to beginners.
You can use the same software on both smartphones and PCs. iPhone support began in April 2026, so you can use Aqua Voice exclusively on your personal iPhone and your home PC. Once you learn how to use the microphone, the experience remains the same across devices, which significantly lowers the initial learning cost.
With the latest update, Japanese conversion has become significantly faster and more accurate. Since supporting the new model (Avalon 1.5), the speed of Japanese conversion has felt completely different, and the accuracy has also improved considerably. Teachers who are still stuck with the old image that 'Japanese is slow and has many misconversions' are worth trying it again.
You can sign up for a personal subscription for $8 a month. You don't need to submit a proposal to install it on your school PC or apply to the IT department. You can install it on your personal smartphone or PC with your own credit card and start trying it tomorrow.
I repeat, the main point is to 'choose the software that suits you.' However, the time spent worrying about the first one is the biggest waste. If you're lost, I tell people to start with Aqua Voice and compare it with other software after you've gotten used to it.
The final form for serious expansion—Shokz OpenComm2 UC × Aqua Voice at home
Once you've created an entry point with EarPods and your workflow is running, it's time to think about the next stage.
Researching teaching materials at home, drafting structures between chores, creating school documents on holidays, and finishing up reports at home. The final form of this battlefield is the combination of Shokz OpenComm2 UC (2025 Upgrade) ¥27,880 (Shokz Official) × Aqua Voice ($8/month when billed annually).
Aqua Voice for your home PC, and Shokz OpenComm2 UC for your home microphone. Only when you have these two can a teacher's writing tasks be 'almost entirely handled just by speaking.' No matter how much you polish Gemini + EarPods on your school PC, this is where you will eventually arrive. Keep this as the map you are aiming for from the start.
Type: Bone conduction wireless
Connection: Bluetooth 5.1
Included: Shokz Loop 120 USB-A/C wireless adapter
Microphone: Boom microphone (with noise canceling and mute button)
Talk time: up to 16 hours / Playback time: up to 8 hours
Charging: USB-C
There are two reasons to choose the 'UC' (2025 Upgrade version), which is 5,000 yen more expensive than the standard OpenComm2. One is the included Loop 120 dongle—when you plug this dongle into a PC, the connection is more stable than standard Bluetooth pairing. The other is that it supports USB-C charging. Being able to charge with the same cable as your PC and smartphone is a small detail, but it makes a difference in long-term use. For those who work on a PC at home for long hours, this 5,000 yen difference is well worth it.
Because it's bone conduction, your ear canals aren't blocked, so you can hear children's voices and the doorbell for deliveries. Dictating the structure of class newsletters during housework, or reciting material for next week's lesson preparation while preparing dinner—this kind of 'while-doing' operation is a perfect match for bone conduction.
To be honest about my own work at home, I use different tools depending on the situation. For tasks that require checking detailed teaching materials or accuracy in proper nouns, I look at the screen and use my hands.
However, when I'm discussing lesson structures or researching interesting ways to teach a unit on the web, note, or YouTube, I keep my Shokz on and talk to Aqua Voice as I proceed. 'Organize the background of this era with an easy-to-understand analogy,' 'I want to write a grade-level newsletter for after Golden Week with this content, please think of a structure'—such consultations and drafting are almost entirely done by dictation. Just by switching from typing to 'speaking and saving,' the density of your preparation changes.
And as a supplement: If you don't mind the microphone sticking out from your ear, you can pair it with your iPhone and use it normally outdoors or on your commute.. You don't have to limit it to home use. If you adjust the angle of the microphone a little, it's barely noticeable at a glance if you have dark hair. If you are okay with the visual and fit, you can cover all battlefields except for your school work PC with this one device.
This is the 'strongest configuration' a teacher can put together at the moment. However, you don't need to aim for this from the start. Create an entry point with EarPods, and keep this as a destination for six months or a year later once your operations are running smoothly.
Small operational rules to lower cognitive load
Stumbling with voice input is often due to 'not knowing what to say' rather than the equipment. Here are three operational techniques to lower the cognitive load at the start.
Numbering operation: When dictating for multiple people, such as for observations, separate them by numbers like 'first person,' 'second person.' If you assign the proper names in the order of the roster after AI formatting, your mouth won't stop even for 35 people.
Scene memo operation: Dictate only the scene briefly, such as 'During lunch preparation...' or 'Regarding the trouble during cleaning...' and accumulate memos -> have the AI generate evaluation words or observation sentences later.
Break design during continuous operation: Don't try to dictate for 35 people in one session. Take a 3-minute break every 10 people. For the sake of your throat and to check the AI's formatting quality, breaking it up is faster in the end.
Combined with the note at the beginning, this is all you need for operational judgment. You can run it as part of your everyday work.
A step from tomorrow—a gradient from individual to grade level to school-wide
Finally, here are the concrete steps for tomorrow onwards.
Today
Apple official siteOrder EarPods (USB-C) for ¥2,780
Open Gemini with your school Google account and press the microphone button once to see if it appears.
Before the EarPods arrive, try talking to Gemini once with the PC's built-in microphone (dictate a draft of a short contact message for 30 seconds -> give formatting instructions). If you experience the process and accuracy, you will notice the difference on the day they arrive.
This week
Once your EarPods arrive, plug them into your school computer and confirm they are recognized as a 'USB Audio Device'
During a quiet time early in the morning or after school, tell a colleague sitting nearby, 'I'm going to try voice input,' and then try writing one student report or one parent newsletter using voice input
This month
During the 10-minute or 20-minute breaks, try asking Gemini once a day via voice input, 'List three points where students might struggle in the next lesson' (A minimal routine for utilizing spare time)
If you have days where you take the train to work, try the workflow of drafting on your iPhone with EarPods and sending it to your school email once (after checking your school's regulations. Those who want to read more are welcome to read the book)
Once you get used to the Gemini voice button, try handing it to one trusted person in your grade level, saying, 'You can try it out for 3,000 yen'
When you start working from home more, install the trial version of Aqua Voice on your home PC to experience the difference in accuracy compared to Gemini
Once you can feel the benefits of Aqua Voice, move on to the final configuration of Shokz OpenComm2 UC × Aqua Voice
To your grade level and school
The right approach is to introduce it modestly. There is no need to spread it to everyone at once. Once you see that your own reports and newsletters are taking several hours less per month, that is enough reason to share it with those around you.
Appendix: Template for application to the IT department (only for schools that require it)
Please use this only for schools that require an application to use a USB audio device. It can serve as material for cases where the IT staff cannot answer immediately or for municipalities that require an inquiry to the Board of Education.
件名:校務PCでのUSBオーディオデバイス使用申請
使用機器:Apple EarPods (USB-C) ¥2,780
ベンダー:Apple Inc.
クラス:USB Audio Class(オーディオ入出力のみ・ストレージ機能なし)
使用目的:Geminiチャット機能の音声入力による
業務文書(通信・所見素案等)作成効率化
使用場所・時間:職員室自席・勤務時間内
個人情報の扱い:固有名は声に出さず番号・場面で口述する運用とするConclusion—Just by changing the microphone, both Aqua Voice and Gemini become something else entirely
¥2,780.
Whether the software is Aqua Voice, Gemini used on a school PC, or the default voice input on an iPhone, the moment you bring the microphone to your mouth, the input experience changes into something else entirely. Cover it with your hand and speak softly in the staff room, draft on your iPhone while commuting, and expand to the final form of Shokz × Aqua Voice at home.
You don't need special permissions, monthly subscriptions, or long approval processes. Just by pressing the order button on the official Apple site, you can start tomorrow.
The core of voice input is the microphone. Just by changing the microphone, the gateway to AI becomes something else entirely.
I would be happy if even one more teacher in the staff room felt that way.

