LITF-PA-2026-183 · Smart Home / Seismology

System and Method for Earthquake Early Warning Using Reverse-Electromechanical Transduction in Networked Smart Speaker Drivers

Smart speaker on a living room shelf with concentric seismic wave rings radiating outward and a glowing P-wave arrival trace

Defensive Prior Art Disclosure. This document describes an invention for the sole purpose of establishing that the invention was publicly known as of the publication date above. It is not a patent application, and no patent rights are claimed. The authors dedicate this disclosure to the public domain to prevent future patenting of the described invention.

Abstract

Earthquake early warning works because compressional P-waves outrun the destructive shear S-waves by seconds to tens of seconds, but only where dense seismic instrumentation exists. Dedicated networks are expensive: the USGS ShakeAlert buildout targets 1,675 stations across three states. This disclosure describes a system that converts the existing installed base of smart speakers into a building-coupled seismic sensor network at zero marginal hardware cost. A moving-coil loudspeaker driver is a reversible electromechanical transducer: when the driver frame vibrates with building motion and the cone lags on its suspension, the voice coil generates a terminal voltage proportional to relative cone velocity (e = Bl · v). During idle windows the speaker amplifier is placed in a high-impedance sensing state and the driver terminals are sampled as a seismic channel; during playback, motion-induced back-EMF is extracted from measured drive current against an electrical model of the driver. An on-device detector (STA/LTA trigger plus a quantized 1D convolutional classifier) distinguishes P-wave onsets from domestic low-frequency confounders such as door slams, subwoofer bass, footfalls, and HVAC transients. Candidate detections with synchronized timestamps are reported to a regional correlator that confirms events on multi-device coincidence consistent with P-wave travel times, estimates the epicenter by time-difference-of-arrival, and issues S-wave arrival warnings to devices in the projected shaking zone. Only sub-20 Hz band features leave the device, so no intelligible audio is transmitted. The result is a hyperlocal early warning layer whose urban sensor density can exceed dedicated networks by orders of magnitude wherever smart speakers are already deployed.

Technical Field

This disclosure relates to earthquake early warning, low-frequency vibration sensing with electrodynamic transducers, on-device machine-learning classification of seismic signals, distributed sensor networks, and privacy-preserving crowdsourced environmental monitoring using consumer smart-home devices.

Background

Earthquake early warning (EEW) exploits a physical head start: P-waves travel through continental crust at roughly 6 km/s while the damaging S-waves and surface waves travel at roughly 3.5 km/s and slower. Sensors near the epicenter detect the P-wave, a processing center estimates location and magnitude, and alerts reach populations farther away seconds before strong shaking arrives. The USGS ShakeAlert system began public alerting in California in 2019 and expanded to Oregon and Washington in 2021, with an implementation plan calling for 1,675 seismic stations; the Pacific Northwest buildout alone recently completed 569 stations across Washington and Oregon. The Congressional Research Service describes the architecture plainly: sensors detect P-waves, data centers process them, and public and private pathways convert alert messages into warnings ahead of the later-arriving, more damaging S-waves.

Dedicated stations are sensitive and well-calibrated, but they are sparse outside high-priority corridors and expensive to install and maintain. Two consumer-hardware approaches have tried to fill the gap. UC Berkeley's MyShake project turns smartphones into a citizen-science seismic network using built-in accelerometers, demonstrating that commodity sensors plus network correlation can contribute detections where traditional networks are thin. Separately, hobbyists have long built seismometers from loudspeaker drivers: with the basket fixed to the ground and added mass on the cone, frame motion relative to the inertial cone generates a measurable voltage, essentially an audio signal fed to a sound card. These DIY instruments prove the transduction principle but remain single-device curiosities, not networks.

A third, overlooked fact completes the picture. The loudspeaker-to-microphone direction has been demonstrated in commodity hardware: Guri et al. (2016, "SPEAKE(a)R") showed that headphones and some loudspeakers connected to a PC can be software-retasked from output to input, capturing intelligible audio from up to 20 feet away through the same reversible transduction. If a driver can serve as a microphone for kilohertz speech, it can serve as a velocity sensor for sub-20 Hz ground motion, provided the sensing path is designed for the seismic band rather than the audio band.

Meanwhile the installed base has quietly become enormous. Edison Research's Infinite Dial 2025 study puts US smart speaker ownership at 35 percent of Americans aged 12 and older, roughly 101 million people; Parks Associates reports 51 percent of US internet households own a smart speaker or smart display. These devices sit on shelves and counters, coupled through furniture and floors to building structures, idle most of the day, each containing a moving-coil driver, a low-noise audio signal path, an applications processor with ML acceleration, a network connection, and a clock. No EEW proposal has combined these facts: the driver as a building-coupled seismometer, the idle audio path as the digitizer, on-device ML as the discriminator, and the installed base as the network.

What is needed is a complete system design that turns this latent network into a working early warning layer: a sensing mode for the driver, an honest accounting of its sensitivity limits, an on-device classifier that rejects domestic confounders, a multi-device correlation scheme that converts weak single-device signals into reliable regional detections, and a privacy design that keeps audio content on the device. The non-obvious step is not the transduction itself, which hobbyists demonstrated with single instruments, but the combination that makes it deployable at scale: recovering a seismic channel during active playback through back-EMF modeling, corroborating it against the microphone array to reject electrical artifacts, and weighting thousands of uncalibrated household devices by learned trust inside a travel-time correlator, so that individually weak sensors compose into a regional detector.

Detailed Description

1. Transduction Principle and Sensitivity

A moving-coil loudspeaker driver consists of a voice coil in a permanent-magnet gap, attached to a cone suspended by a surround and spider with combined compliance C and moving mass m, giving a free-air resonance fs = 1/(2π√(mC)), typically 70 to 100 Hz for the small drivers in smart speakers. Driven electrically, coil current in the magnetic field produces force (F = Bl · i). In reverse, relative velocity v between coil and magnet generates an open-circuit terminal voltage e = Bl · v, where Bl is the force factor (typically 2 to 5 T·m for small drivers).

For seismic sensing the frame is the moving reference. Ground acceleration a(t) couples through the building, floor, furniture, and enclosure to the driver frame. The cone and coil, possessing inertia, lag the frame; the suspension supplies the restoring force. For excitation frequencies f well below fs, the system is stiffness-controlled and the relative displacement is approximately xrel ≈ a/(2πfs)². Relative velocity at the excitation frequency is vrel = 2πf · xrel, and the generated voltage is e = Bl · vrel.

Illustrative numbers (not measured from any prototype): for strong shaking with peak ground acceleration a = 0.5 m/s² (approximately 0.05 g), a driver with fs = 80 Hz gives xrel ≈ 0.5/(2π·80)² ≈ 2 × 10-6 m. At a representative P-wave frequency f = 2 Hz, vrel ≈ 25 × 10-6 m/s, and with Bl = 3 T·m the terminal voltage is e ≈ 75 µV. A 24-bit audio ADC path with a low-noise preamp resolves microvolt signals with margin; the honest comparison is that single-device sensitivity sits roughly 40 to 60 dB below a dedicated geophone. The system compensates with network density rather than single-device sensitivity, and detection thresholds are set for moderate-to-strong local shaking (approximately magnitude 4.5 and above within roughly 100 km), not teleseismic recording.

2. Sensing Path Embodiments

Embodiment A: idle-window sensing. Smart speakers spend most of the day not playing audio. When playback is inactive, the audio amplifier is placed in a high-impedance state (or the driver is routed through an analog multiplexer) and the driver terminals are connected to a low-noise instrumentation preamp feeding the existing audio ADC at 200 Hz or higher sample rate. A digital lowpass filter with a corner near 25 Hz followed by decimation to 100 Hz preserves the seismic band (0.3 to 20 Hz) while rejecting audio-band interference. Duty cycle is inherently high because idle dominates the usage profile; sensing pauses only during active playback and voice-assistant interactions.

Embodiment B: back-EMF extraction during playback. Many smart speakers already measure driver current for speaker-protection functions (thermal limiting and excursion prediction). During playback the terminal voltage and current are both known; subtracting the driver's electrical model (DC resistance Re and voice-coil inductance Le, calibrated per device) from the measured terminal voltage recovers the motion-induced back-EMF, which contains the seismic component superimposed on the audio. Signal-to-noise is poorer than in Embodiment A because the audio signal dominates, but adaptive cancellation of the known drive signal yields a usable seismic channel with no sensing gaps.

Embodiment C: microphone-array corroboration. Strong shaking couples into room air as low-frequency pressure fluctuations that the device's microphone array captures. The mic channel is band-limited to the same 0.3 to 20 Hz range and used only as a corroborating feature stream fused with the driver-EMF channel. Agreement between the electromechanical channel and the acoustic channel suppresses false triggers from purely electrical artifacts (amplifier glitches, power-supply transients) that cannot appear in both domains simultaneously.

3. On-Device Detection and Classification

Detection runs entirely on the device's existing applications processor or ML accelerator. A first stage applies a short-term-average/long-term-average (STA/LTA) trigger, the classic seismic onset detector, to the bandpassed driver-EMF signal with a short window of 0.5 s and a long window of 10 s. STA/LTA crossings produce candidate onsets with low compute cost.

A second stage classifies each candidate with a quantized one-dimensional convolutional neural network (four convolutional layers, tens of thousands of parameters, under 100 KB, inference under 5 ms on a typical smart-speaker SoC). The classifier input is a 10-second window of the three available channels (driver EMF, mic-array low band, and, where available, the device accelerometer if the platform exposes one). Output classes include P-wave onset, S-wave/coda energy, and a confounder set: door slam, subwoofer or music bass, footfalls, HVAC transient, appliance vibration (washer spin, refrigerator compressor start), nearby construction or blasting, and freight-train passby. Training data combines recorded shake-table reproductions of earthquake waveforms through representative drivers, field recordings of confounders in real homes, and synthetic augmentation (time-stretching, amplitude scaling, and convolution with measured building transfer functions). The classifier emits a calibrated onset probability; only candidates above a per-device adaptive threshold generate network reports.

Crucially, classification and feature extraction happen on the device. The uplink payload for a candidate detection contains only: a timestamp, a device pseudonym, coarse geolocation (neighborhood granularity), the onset probability, peak band-limited amplitude, dominant frequency, and P-versus-S classification. No waveform audio above 20 Hz and no intelligible audio content ever leaves the device; the seismic band (0.3 to 20 Hz) sits entirely below the fundamental frequency of human speech (approximately 85 Hz), so reconstruction of conversations from the uplink is physically impossible.

4. Network Correlation, Location, and Warning

Candidate detections flow to a regional correlator operated by the platform or a partner alerting authority. A single-device trigger is never enough: confirmation requires at least three devices within a travel-time-consistent radius reporting onsets whose inter-arrival times are consistent with P-wave propagation at approximately 6 km/s from a common source. This coincidence requirement converts individually weak, noisy sensors into a reliable network detector, directly analogous to how ShakeAlert requires four stations before issuing messages.

On confirmation, the correlator estimates the epicenter by time-difference-of-arrival across the detecting devices and estimates magnitude from the early amplitude growth rate and the spatial extent of detections. Each detecting device's P-to-S interval provides an independent hypocentral distance estimate that refines the solution. Devices in the projected shaking zone, including those that have not yet detected anything, receive warning messages keyed to predicted S-wave arrival time: typically seconds for locations near the epicenter, tens of seconds at regional distances. Warnings are delivered through the speaker's own audio output (a distinct alert tone followed by a short spoken instruction), through companion phone push notifications, and through smart-home integrations (pausing stoves, opening smart locks for egress where the user has opted in). The design explicitly provides a feed interface for authorized alerting authorities so detections can supplement systems like ShakeAlert and dissemination can use official channels such as Wireless Emergency Alerts, rather than the platform issuing freelance public alerts.

The density argument is the core of the invention. With roughly 101 million Americans owning smart speakers and about half of US internet households owning a speaker or display, even a low single-digit opt-in percentage yields sensor counts in a large metro that exceed dedicated networks by orders of magnitude: a 2 percent opt-in across California's roughly 13 million households is about 260,000 sensors, versus 1,675 planned ShakeAlert stations across three states. Urban faults, where people and shaking coincide, are exactly where speaker density is highest.

5. Calibration and Enrollment

Each device self-calibrates on enrollment and periodically thereafter. An inaudible low-frequency probe (a swept sine from 5 to 120 Hz at low level, or analysis of the driver's response to known playback content) estimates fs and Bl, setting the expected sensitivity scale. The classifier threshold adapts per device over a two-week burn-in, learning the household's confounder profile (proximity to rail lines, subwoofer ownership, construction schedules). Users opt in explicitly; enrollment explains in plain terms what is sensed (sub-audible vibration only), what leaves the device (timestamps and band-limited features), and what the warnings mean. A quiet self-test mode lets users verify the warning path without a real event.

Claims

  1. A seismic sensing system comprising: a smart speaker having a moving-coil loudspeaker driver with a frame, a magnet assembly, and a voice coil suspended on a compliant suspension; a sensing circuit configured to capture, as a seismic signal, the terminal voltage generated by relative motion between the voice coil and the magnet assembly when building motion vibrates the frame; and a processor configured to detect earthquake P-wave onsets in the seismic signal.
  2. The system of claim 1, wherein the sensing circuit comprises a high-impedance sensing state engaged during audio-playback idle windows, routing driver terminal voltage to an analog-to-digital converter sampling at 200 Hz or higher, followed by digital lowpass filtering preserving a seismic band of approximately 0.3 to 20 Hz.
  3. The system of claim 1, wherein the sensing circuit measures driver terminal voltage and current during audio playback and subtracts a calibrated electrical model of the driver to recover motion-induced back-EMF as the seismic signal.
  4. The system of claim 1, further comprising a microphone array whose output is band-limited to the seismic band and fused with the driver seismic signal as a corroborating channel, wherein agreement between the electromechanical channel and the acoustic channel suppresses electrically induced false triggers.
  5. The system of claim 1, wherein the processor applies a short-term-average/long-term-average trigger followed by a quantized on-device neural network classifier that distinguishes P-wave onsets from a confounder set comprising door slams, music bass, footfalls, HVAC transients, appliance vibration, construction blasting, and train passbys.
  6. The system of claim 1, wherein the processor transmits to a network correlator only sub-20 Hz band features, onset timestamps, and coarse geolocation, and wherein no intelligible audio content leaves the device.
  7. A regional earthquake detection network comprising a plurality of systems according to claim 1 and a correlator configured to confirm an earthquake event when at least three systems report onsets with inter-arrival times consistent with P-wave propagation from a common source, and to estimate an epicenter by time-difference-of-arrival across the reporting systems.
  8. The network of claim 7, wherein the correlator estimates hypocentral distance at each reporting system from its P-to-S wave interval and issues S-wave arrival warnings to systems in a projected shaking zone prior to S-wave arrival.
  9. The network of claim 7, wherein warnings are delivered through the smart speaker's own audio output as an alert tone plus spoken instruction, and through at least one of a companion-device push notification and a smart-home automation action.
  10. The network of claim 7, further comprising an interface exporting confirmed detections as a supplemental sensor feed to an authorized earthquake early warning system and using official alert dissemination channels.
  11. The system of claim 1, further comprising a self-calibration routine that estimates the driver's free-air resonance frequency and force factor from a low-frequency probe signal and adapts detection thresholds to a per-household confounder profile learned during a burn-in period.
  12. A method of earthquake early warning without dedicated seismic instrumentation, comprising: operating a moving-coil driver of a networked smart speaker as a velocity transducer for building-coupled ground motion; digitizing driver terminal voltage in a seismic band of approximately 0.3 to 20 Hz; classifying P-wave onsets on the device while rejecting domestic confounders; reporting sub-20 Hz features and timestamps to a regional correlator; confirming events on multi-device P-wave travel-time coincidence; and disseminating S-wave arrival warnings to devices in the projected shaking zone.
  13. The network of claim 7, wherein the correlator maintains a reputation weight for each reporting system based on its historical agreement with confirmed events, and wherein reports from systems with low reputation weights are discounted or excluded from event confirmation.

Implementation Notes

This disclosure describes a proposed design; no prototype has been built and no performance figures have been measured. Several practical considerations apply. Single-device sensitivity is roughly 40 to 60 dB below a dedicated geophone in system-level detection terms, so the system is designed for moderate-to-strong local shaking (approximately magnitude 4.5 and above within roughly 100 km), not for teleseismic recording or precise magnitude estimation from single stations. Building coupling varies enormously: a speaker on a rigid ground-floor slab couples well, while one on a soft couch or an upper floor of a flexible high-rise sees amplified, filtered, or delayed motion, which is why per-device calibration and network coincidence are load-bearing parts of the design rather than refinements. Idle-window sensing misses events during loud playback; the back-EMF embodiment narrows but does not eliminate this gap. Freight trains, quarry blasts, sonic booms, and pile drivers produce legitimate low-frequency energy that the classifier must reject, and adversarial or faulty devices reporting bogus onsets must be downweighted by the correlator's reputation tracking. Public alerting carries regulatory and liability obligations: the intended path is supplementation of authorized systems and use of official dissemination channels, not independent public sirens. Finally, user trust depends on the privacy design holding under scrutiny: all classification on-device, band-limited uplink below the speech range, explicit opt-in, and no waveform retention beyond the detection window.

Prior Art References

  1. Guri, M., Solewicz, Y., Daidakulov, A., and Elovici, Y., "SPEAKE(a)R: Turn Speakers to Microphones for Fun and Profit," arXiv:1611.07350 [cs.CR], 2016, demonstrating software-only retasking of headphone and loudspeaker outputs as inputs and establishing the reversibility of commodity electrodynamic transducers. arxiv.org
  2. Congressional Research Service, "The ShakeAlert Earthquake Early Warning System," IF12956, describing P-wave detection by sensor networks, data-center processing, and alerting ahead of later-arriving damaging S-waves; public alerting in California from 2019 and Oregon/Washington from 2021. congress.gov
  3. Pacific Northwest Seismic Network / phys.org, "With ShakeAlert installations complete, researchers explore offshore expansion," June 2026, reporting 569 completed stations across Washington and Oregon within the 1,675-station ShakeAlert implementation plan. phys.org
  4. UC Berkeley Seismological Laboratory, MyShake Earthquake Alerts, a smartphone-based global seismic network in which phones act as mini-seismometers contributing to earthquake detection. appbrain.com
  5. AutoSpeed, "The DIY Seismometer," documenting construction of a seismometer from a loudspeaker driver with the basket fixed and cone inertia generating a voltage on vibration. autospeed.com
  6. Physics Forums, "What is the physics behind seismometers?", discussion of the speaker-plus-mass seismometer design and differentiation of P-wave and S-wave motion signatures. physicsforums.com
  7. GlobalSpec, "Low-cost earthquake early warning device may soon be available to consumers and businesses," June 2021, covering a sub-$100 consumer geophone-based early warning device filling the gap left by station-based systems. insights.globalspec.com
  8. Edison Research, Infinite Dial 2025, via the secondary compilation at 16best.net: 35 percent of Americans aged 12 and older (about 101 million people) own a smart speaker. 16best.net
  9. Parks Associates, via the secondary compilation at XtendedView, 2025: 51 percent of US internet households report owning a smart speaker and/or smart display. xtendedview.com
  10. Given, D.D., et al., "ShakeAlert earthquake warning: The challenge of transforming ground motion into protective actions," USGS Publications Warehouse, 2021, on the operational system, station requirements (four stations before message issuance), and machine-to-machine protective actions. pubs.usgs.gov