Apple Details S11 Security Architecture Behind New Apple Watch 'Audio Intelligence' Features
The Apple Watch Series 12 and Ultra 4 introduce rolling ambient audio recording, speech summarization, and local memory enclave protections.

Apple announced on Wednesday that its upcoming Apple Watch Series 12 and Ultra 4 smartwatches will feature four new opt-in "audio intelligence" capabilities powered by the devices' built-in microphones, according to a company technical report first reported by Wired. The features include automated environmental sound detection, song identification, conversation summarization, and a "Live Rewind" transcription tool that allows users to view a text conversion of ambient speech recorded over the preceding 15 seconds.
Seeking to address privacy concerns associated with persistent microphone monitoring on wearable devices, Apple emphasized that all four features process data locally where possible and prohibit raw audio files from being accessed by the watch operating system, third-party software applications, end users, or Apple itself. The system architecture relies on the company's new S11 silicon chip, which incorporates a dedicated hardware-isolated memory partition termed the "Secure Exclave" designed specifically for handling sensitive sensor data.
For localized functions such as Sound Recognition, the wearable processes audio entirely on-device to alert users to environmental cues—including doorbells, sirens, emergency alarms, and crying infants—without transmitting data off the watch. Meanwhile, an updated integration with Apple-owned music service Shazam generates an encrypted song signature inside the Secure Exclave when ambient music is detected. The smartwatch sends only the mathematical signature to Shazam's servers for identification rather than an audio recording, and the signature is immediately erased from the device once matched or dismissed.
The new conversation logging tool, named Siri Recap, can operate continuously or follow a user-defined schedule. A background artificial intelligence model monitors for speech without keeping continuous audio files or transcriptions. Upon detecting a conversation, the watch encrypts the captured audio within the Secure Exclave buffer and transfers the file to the Secure Exclave of a paired iPhone over an encrypted Bluetooth connection before instantly purging the source audio from the smartwatch.
The paired iPhone processes the incoming audio using on-device speech recognition and language models to produce a streamlined transcript free of extra words, redundant phrases, and potentially harmful language. The raw recording is immediately purged from the iPhone, and a screening safety model evaluates the distilled text. The iPhone then encrypts the text and transmits it to Apple's Private Cloud Compute infrastructure for final processing by the company's foundation models.
To enrich summary quality, Siri Recap attaches non-precise contextual metadata to the request sent to Private Cloud Compute. This contextual package can include active media details from Now Playing, calendar entries to generate accurate meeting titles, and high-level location designations such as home, work, school, or general points of interest like a grocery store or park. Precise geographic coordinates and specific business names are excluded, and Apple's cloud models are programmed to automatically redact sensitive personal details, including financial account data and government identification numbers, before returning an encrypted summary to the user's devices.
The "Live Rewind" feature utilizes the S11 chip's Secure Exclave to maintain a continuous, rolling 15-second buffer of environmental audio, constantly overwriting older data. Users initiate a transcription by double-pressing the Apple Watch's Digital Crown, which prompts the watch to transfer the 15-second buffer to the paired iPhone for local speech-to-text conversion. Once transcribed, the raw audio is erased from the phone, and the text is sent back to the smartwatch screen, where the user can save it to the Siri app or discard it.
To prevent covert audio transcription of surrounding individuals, Apple built hardware and software safeguards into the system. Activating Live Rewind triggers an audible chime from the Apple Watch to alert nearby people that a transcript is being generated, an alert that sounds even if the watch is set to silent mode or connected to wireless headphones. Additionally, both Live Rewind and Siri Recap contain automatic fail-safes that immediately erase stored audio buffers if the secure transfer to a paired iPhone fails or is interrupted.
The new audio intelligence tools reflect Apple's expanding effort to bring generative artificial intelligence and natural language processing to its wearable product line. While the tech giant has built isolated silicon buffers and private cloud protocols to minimize data exposure, the rollout highlights how ambient machine learning capabilities are expanding across consumer hardware platforms.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



