System and Method for Open-Window and Open-Door Detection via Active Acoustic Room-Response Probing on Consumer Smart Speakers with HVAC Energy-Waste Estimation
Abstract
Disclosed is a system and method that converts consumer smart speakers into open-window and open-door detectors with no additional hardware. Each speaker periodically emits a short, low-level near-ultrasonic probe chirp during quiet moments and records the room's acoustic response on its microphone array. Deconvolution of the recorded signal against the known probe yields the room impulse response, from which sub-band reverberation time, early decay time, and early-reflection structure are extracted. An open window or door changes the room's boundary conditions: the opening behaves as a near-perfect absorber, reverberation time falls, specific early reflections weaken or vanish, and the spectral pattern of the decay shifts. A per-room adaptive baseline, conditioned on HVAC fan state and compensated for temperature-dependent sound speed, feeds a change-point detector. The magnitude and spatial pattern of the change, resolved by microphone-array beamforming, classify the event as a window, an exterior door, or an interior door, with multi-speaker corroboration across rooms. Correlation with thermostat state estimates the energy wasted while conditioning against an open envelope, expressed in kilowatt-hours and dollars, and generates user nudges. All processing runs on-device; raw audio is discarded immediately and only scalar acoustic features leave the device.
Field of the Invention
This invention relates to smart-home sensing, specifically to detecting open windows and doors by active acoustic probing of room impulse responses using the loudspeakers and microphone arrays of consumer smart speakers, and to estimating HVAC energy waste attributable to conditioning against an open building envelope.
Background
Space heating and air conditioning are the largest end uses of energy in American homes, according to the U.S. Energy Information Administration's Residential Energy Consumption Survey. Running an HVAC system while a window or door stands open is pure waste: conditioned air escapes directly, the equipment runs longer duty cycles, and in humid climates the system also fights a continuous latent load it was never sized for. Yet most homes have no way to notice. The person who cracked the bedroom window on a mild evening forgets it when the night turns cold and the furnace kicks on.
Existing detection approaches all carry adoption friction. Per-opening contact sensors require a sensor, a battery, and a pairing step for every window and door; a typical house has twenty or more openings, and few households instrument all of them. Thermostat-based inference watches for temperature anomalies, but it is slow (tens of minutes) and confounded by solar gain, cooking, and occupancy. Cameras can see an open window but raise privacy objections in bedrooms and cost far more per room. What every room already contains, in tens of millions of homes, is a smart speaker: a calibrated loudspeaker paired with a microphone array and a capable processor.
Room acoustics gives the speaker something to measure. A room's impulse response, the sound field that results from an idealized instantaneous excitation, encodes the geometry and absorption of every boundary. Standard practice measures it with logarithmic sine sweeps or maximum-length sequences and derives reverberation time by Schroeder backward integration of the squared response (Schroeder, J. Acoust. Soc. Am., 1965; ISO 3382-1:2009). An opening in the envelope is, acoustically, a patch of near-unit absorption: sound that would have reflected now leaves the room. The effect is strongest at high frequencies, where wavelengths are centimeters and even a modest window is many wavelengths across, so reverberation time in the upper bands falls measurably and the early reflections that used to arrive from that wall weaken or disappear. Architectural acoustics textbooks treat this boundary-absorption accounting as elementary (Kuttruff, Room Acoustics, 6th ed., 2016).
Commodity speaker-microphone pairs have already been repurposed as active acoustic sensors. Nandakumar, Gollakota, and Sunshine demonstrated sonar-like breathing and sleep-apnea detection using inaudible probes emitted by an ordinary smartphone speaker (Proc. ACM MobiSys, 2015), and Gupta et al. used the Doppler shift of a device's own speaker output for gesture sensing (SoundWave, Proc. ACM CHI, 2012). These systems sense physiology and gesture. They do not probe the room itself, do not track reverberation parameters over time, and do not connect acoustic observations to the building envelope or to HVAC operation.
Device-free acoustic sensing research has used channel variation for occupancy and activity inference, but that work is passive and concerned with people, not with the envelope. The sibling disclosure LITF-PA-2026-168 applies acoustic pulse reflectometry to duct leakage: a related technique aimed at a different target (ductwork rather than rooms) with a different measurand (reflections in a one-dimensional waveguide rather than three-dimensional room reverberation).
The gap is therefore: active, inaudible room-response probing on hardware already in the room, with an adaptive baseline that learns each room's normal acoustic fingerprint, classification of the opening by type, corroboration across multiple speakers, and a direct coupling to thermostat state that turns an acoustic observation into a quantified energy-waste figure the homeowner can act on.
Detailed Description
1. Probe signal design
Each speaker emits a probe consisting of a logarithmic sine sweep over a near-ultrasonic band, disclosed as an example band of 18 to 22 kHz with a duration of 150 to 300 milliseconds. The band sits above the hearing range of most adults while remaining below the 24 kHz Nyquist limit of the 48 kHz sampling used by commodity smart-speaker audio pipelines, so no hardware modification is required. The sound pressure level is kept low, disclosed as an example value of approximately 40 dBA at 1 meter, which is below typical room background in the voice band and contributes negligibly to the ultrasonic noise floor.
Probes are emitted only during quiet windows: no media playback, no voice-assistant session in progress, and a voice-activity detector reporting silence for a preceding guard interval, disclosed as an example value of 10 seconds. The interval between probes is randomized within a disclosed example range of 10 to 30 minutes, which prevents entrainment with periodic household sounds and keeps the average duty cycle negligible. A household with small children or pets may select a higher probe band, disclosed as an example of 20 to 23 kHz at reduced level, since children can hear 18 kHz tones and dogs hear well into the ultrasonic range; the band is a configuration parameter, not a fixed constant.
2. Room impulse response estimation and acoustic features
The microphone array records the probe, and the system deconvolves the recording with the known sweep (inverse filtering) to recover the room impulse response as seen from the speaker's position. A loopback calibration step normalizes for the individual speaker's transducer response: the direct-path arrival, the first and strongest peak, serves as the reference, so differences between speaker models in tweeter and microphone frequency response divide out.
From the impulse response the system computes, per probe, a feature vector: third-octave-band reverberation times RT60 in the probe band (via Schroeder backward integration of the band-filtered squared response), early decay time (EDT, from the first 10 dB of decay), clarity C50, center time, and the envelope of the first 50 milliseconds, which encodes the early-reflection structure. These quantities are standard room-acoustic parameters (ISO 3382-1:2009); the novelty lies in tracking them continuously on a consumer device rather than measuring them once with laboratory equipment.
3. Adaptive baseline and environmental compensation
Every room has a different acoustic fingerprint, so the system learns a per-room baseline over a disclosed example period of 7 to 14 days, then tracks it with an exponential moving average gated on quiescent state (no event declared, no large transient). The baseline is conditioned on HVAC fan state, obtained from the thermostat, because forced-air operation raises the background noise floor and slightly alters the acoustic field through duct paths; probes are preferentially scheduled for fan-off intervals, and fan-on probes are compared against the fan-on baseline.
Temperature compensation corrects for the temperature dependence of the speed of sound, c = 331.3 × √(1 + T/273.15) meters per second: a 1 °C change shifts arrival times by a fraction that would otherwise masquerade as a geometric change. The system reads the thermostat's temperature sensor (or its own, where fitted) and time-warps the measured response to a reference temperature before feature extraction. Humidity enters at second order and is folded into the baseline adaptation rather than modeled explicitly. A gross baseline shift, such as the speaker being moved to a new shelf, is detected as a persistent large deviation and triggers re-learning with a user confirmation prompt rather than a stream of false events.
4. Detection, classification, and multi-speaker fusion
Detection is change-point detection on the feature vector: the Mahalanobis distance from the current measurement to the adaptive baseline, using the baseline's learned covariance, is compared against a threshold, and an event is declared when the distance exceeds the threshold for a disclosed example value of three consecutive probes. Requiring persistence rejects transients such as a person walking through the room, which disturb the field for one probe interval and then resolve.
Classification uses the magnitude, spectral pattern, and spatial signature of the change. A window opening produces a large drop in high-band RT60 with a directional loss of early reflections: microphone-array beamforming identifies which wall's reflections weakened, and the wall bearing the room's windows is known from a one-time setup (or inferred from the largest historical change). An exterior door left ajar produces a still larger change and is corroborated by outdoor noise ingress, a rise in the voice-band noise floor that the probe itself does not measure but the microphones observe continuously. An interior door opening produces a smaller, symmetric change: the two coupled rooms share the opening, so it is corroborated by a complementary change reported by the speaker in the adjacent room. The principal confounder, curtains drawn across a window, adds high-frequency absorption without removing the reflection path and without outdoor noise coupling; where ambiguity remains, the system asks the user once ("was that the curtains?") and records the labeled example for that room.
Multi-speaker fusion runs at the feature level with timestamps; no raw audio crosses between devices. Each speaker reports its event hypotheses with confidence, and a home hub (or a designated primary speaker) applies the corroboration logic: complementary changes in adjacent rooms confirm an interior door, an isolated change with noise ingress confirms an exterior opening, and a change seen weakly everywhere suggests a whole-house state change such as the HVAC fan rather than any single opening.
5. HVAC correlation and energy-waste estimation
The acoustic event becomes actionable when joined to thermostat state. When an opening event coincides with an active call for heating or cooling, the system estimates the excess energy by comparing the HVAC duty cycle during the event against the baseline duty cycle for matched outdoor conditions, binned by outdoor temperature in disclosed example bins of ±1 °C, using recent history or a weather-service feed. The duty-cycle delta multiplied by the equipment's power draw (nameplate rating, or a learned value from a smart thermostat's runtime reports) and the event duration yields kilowatt-hours; multiplied by the user's utility rate, it yields dollars.
The user-facing output is a nudge, not an alarm: "The bedroom window has been open for 40 minutes with the air conditioning running, about $0.60 of cooling so far." Escalation is staged: a first notice, a reminder if the state persists, and, where the user has opted in, a suggested thermostat setback. When the opening closes, the system detects the acoustic return to baseline, ends the event, and reports the total. Aggregated over a month, the system reports the household's open-envelope waste as a single figure, which is the quantity that changes behavior.
6. Privacy and data minimization
The probe band excludes the voice band entirely, and the feature extraction operates only on the deconvolved impulse response in the probe band; no speech content is ever processed, stored, or transmitted as part of this system. All processing runs on the speaker's own processor. Raw microphone audio is discarded immediately after each probe's features are extracted, with no retention. The only data that leaves the device are scalar acoustic features (band RT60 values, event flags, confidence scores) and the waste estimates derived from them, which reveal nothing about conversations, occupancy patterns beyond the disclosed events, or room contents. Enrollment is explicit opt-in, presented as envelope monitoring.
7. Security mode
The same event detector serves a security function. In night mode or away mode, an exterior opening event (window or exterior door, distinguished by the classification of Section 4) generates a security notification rather than an energy nudge: "Front door ajar at 11:42 PM" or "Living room window opened while the house is in away mode." Expected openings are dismissible, and the system learns recurring patterns (the kitchen window opened every evening for cooking) to suppress routine alerts while still flagging anomalies.
8. Figure descriptions
Figure 1 shows the system architecture: smart speakers in multiple rooms emitting near-ultrasonic probes, on-device impulse-response estimation, the adaptive baseline and change-point detector, the thermostat/HVAC interface, the fusion hub, and the user notification path. Figure 2 is a signal diagram: the logarithmic sweep probe, the recorded response, the deconvolved room impulse response, and the Schroeder decay curve with the RT60, EDT, and early-reflection regions marked, shown for a closed-window baseline and an open-window measurement. Figure 3 shows an example event timeline: probe-derived RT60 dropping at the moment a window opens, the thermostat's call-for-cooling state, the duty-cycle comparison against the matched-outdoor-temperature baseline, and the accumulating dollar figure until the window closes and the trace returns to baseline.
Claims
- A system for detecting open windows and doors, comprising: at least one consumer smart speaker having a loudspeaker and a microphone array; a probe module configured to emit a near-ultrasonic probe signal during quiet windows; a room-response module configured to deconvolve the microphone recording against the probe signal to obtain a room impulse response and to extract reverberation features therefrom; a baseline module configured to maintain an adaptive per-room baseline of the reverberation features; a detection module configured to declare an opening event by change-point detection of the features against the baseline and to classify the event as a window, an exterior door, or an interior door; an HVAC interface configured to read thermostat state; and a notification module configured to report the event and an associated energy-waste estimate to a user.
- The system of claim 1, wherein the probe module emits a logarithmic sine sweep confined to a near-ultrasonic band above the hearing range of most adults and below the Nyquist limit of the speaker's audio sampling rate, at a sound pressure level below typical room background, and only when no media playback or voice session is active and a voice-activity detector reports silence for a guard interval.
- The system of claim 1, wherein the room-response module normalizes for transducer variation by using the direct-path arrival of the probe as a loopback reference, computes third-octave-band reverberation times by Schroeder backward integration of the band-filtered squared impulse response, and further extracts early decay time, clarity, center time, and the early-reflection envelope.
- The system of claim 1, wherein the baseline module conditions the adaptive baseline on HVAC fan state obtained from the thermostat, time-warps measured responses to a reference temperature using the temperature dependence of the speed of sound, and triggers re-learning with user confirmation upon detecting a persistent gross deviation indicative of speaker relocation.
- The system of claim 1, wherein the detection module computes the Mahalanobis distance of the feature vector from the baseline covariance and declares an event only after the distance exceeds a threshold for a plurality of consecutive probes, thereby rejecting single-probe transients caused by occupants moving through the room.
- The system of claim 1, wherein the detection module classifies the event using the magnitude and spectral pattern of the reverberation change, microphone-array beamforming that identifies which wall's early reflections weakened, and corroborating outdoor noise ingress observed in the voice band, distinguishing windows from exterior doors and from interior doors.
- The system of claim 6, wherein the detection module discriminates curtains drawn across a window from an open window by the absence of reflection-path removal and the absence of outdoor noise coupling, and records a one-time user label for ambiguous cases in that room.
- The system of claim 1, further comprising a plurality of smart speakers in different rooms and a fusion component that corroborates event hypotheses at the feature level using timestamps without exchanging raw audio, confirming interior-door events by complementary changes in adjacent rooms and attributing whole-house changes to HVAC state rather than to any single opening.
- The system of claim 1, wherein the HVAC interface correlates the opening event with thermostat call-for-heating or call-for-cooling state and estimates wasted energy by comparing HVAC duty cycle during the event against a baseline duty cycle for matched outdoor-temperature bins, converting the duty-cycle delta via equipment power draw and event duration into kilowatt-hours and currency, and wherein the notification module issues a staged nudge reporting the accumulating cost and, upon detecting acoustic return to baseline, reports the event total.
- The system of claim 1, wherein all probe processing runs on the speaker's own processor, the probe band excludes the voice band, raw microphone audio is discarded immediately after feature extraction with no retention, and only scalar acoustic features and derived estimates leave the device.
- The system of claim 1, wherein the notification module operates in a security mode during night or away periods, generating a security notification for exterior opening events classified as windows or exterior doors, suppressing alerts for user-dismissed recurring patterns.
- A method of detecting open windows and doors and estimating associated HVAC energy waste, comprising: emitting, from a consumer smart speaker during a quiet window, a near-ultrasonic probe signal; recording the probe signal on the speaker's microphone array; deconvolving the recording against the probe signal to obtain a room impulse response; extracting reverberation features comprising sub-band reverberation times and early-reflection structure; comparing the features against an adaptive per-room baseline conditioned on HVAC fan state and compensated for temperature-dependent sound speed; declaring and classifying an opening event by persistent change-point detection; correlating the event with thermostat heating or cooling state; estimating wasted energy from HVAC duty-cycle deviation under matched outdoor conditions; and notifying a user with the accumulating cost; wherein raw audio is discarded immediately and no voice-band content is processed.
Implementation Notes
The quiet-window scheduler is the practical heart of the system: a probe emitted over music or conversation is wasted, so the guard interval and the fan-state conditioning do most of the work of keeping the data clean. In practice the richest probing opportunities are the small hours, when the house is silent and the HVAC often cycles; an opening left overnight is also the most expensive kind, so the schedule aligns naturally with the value. Transducer calibration via the direct-path arrival matters more than it first appears: tweeter response at 20 kHz varies by several decibels across speaker models and even across units, and without the loopback normalization the baseline would learn the speaker rather than the room. Open-plan spaces produce weaker effects because there are fewer enclosing boundaries to begin with; the early-reflection envelope carries the detection there rather than the RT60 number. Apartments need the voice-band stability gate, since a neighbor's television through a shared wall looks like noise ingress without an opening. The child/pet probe band is not a courtesy feature but a correctness one: a dog that hears every probe will investigate the speaker, and the resulting near-field disturbance corrupts the measurement far more than the slightly reduced bandwidth costs.
Limitations
The system detects changes in the acoustic boundary condition; it cannot see an opening that produces no measurable change, such as a window opened a crack in an already highly absorptive room. Heavy curtains drawn across a closed window are the principal confounder and are handled heuristically, not perfectly. Very reverberant small rooms (tiled bathrooms) can saturate the measurement, and speakers placed inside cabinets or behind furniture couple poorly to the room. Pets and small children may hear the probe despite the ultrasonic band and low level, which is why the band is configurable. The energy-waste estimate depends on thermostat integration and on outdoor-temperature-matched baselines; without either, the system still detects openings but cannot price them. Privacy guarantees depend on an honest on-device implementation; the protocol minimizes what leaves the device but cannot constrain a malicious client. This disclosure is a monitoring and nudge layer, not a substitute for a security system.
Prior Art References
- ISO 3382-1:2009, Acoustics — Measurement of room acoustic parameters — Part 1: Performance spaces: standard methods for room impulse response measurement using sine sweeps and maximum-length sequences; definitions of RT60, early decay time, and clarity
- Schroeder, M. R., "New Method of Measuring Reverberation Time," Journal of the Acoustical Society of America, 37(3):409–412, 1965: backward integration of the squared impulse response for reverberation-time estimation
- Kuttruff, H., Room Acoustics, 6th ed., CRC Press, 2016: boundary absorption, room modes, and the effect of openings and absorptive patches on reverberation
- Nandakumar, R., Gollakota, S., and Sunshine, J., "Contactless Sleep Apnea Detection on Smartphones," Proc. ACM International Conference on Mobile Systems, Applications, and Services (MobiSys), 2015: inaudible sonar-like probes from a commodity smartphone speaker and microphone for physiological sensing; establishes the speaker-microphone pair as an active acoustic sensor
- Gupta, S., Morris, D., Patel, S. N., and Tan, D., "SoundWave: Using the Doppler Effect to Sense Gestures," Proc. ACM Conference on Human Factors in Computing Systems (CHI), 2012: commodity device speaker and microphone used as a Doppler gesture sensor
- U.S. Energy Information Administration, Residential Energy Consumption Survey (RECS): space heating and air conditioning as the largest end uses of energy in U.S. homes
- Reverberation, Wikipedia: reverberation time definition and its dependence on room volume and boundary absorption
- LITF-PA-2026-168: System and Method for Duct Leakage Detection via Acoustic Pulse Reflectometry Using Smart Speaker Infrastructure, Live in the Future defensive prior art: sibling disclosure applying acoustic reflectometry to ductwork rather than room reverberation