Cap buffered audio and keep transcripts out of INFO logs

Wyoming has no authentication, so buffered audio is attacker-controlled.
self.audio grew until AudioStop with no bound: measured at ~11 MB/s over
loopback, one connection exhausts 32 GB in under an hour, and a stuck
satellite that never sends AudioStop does the same by accident. Cap it at
--max-audio-seconds (default 120), dropping the excess with a single warning
while still transcribing what was captured. Verified: a client streaming
10.8 GB now moves server RSS by 213 MB rather than 10.8 GB.

Transcripts were logged at INFO. Log files are long-lived and world-readable
under /tmp on macOS, so every voice command sat in plaintext readable by any
local account. INFO now records duration, latency and character count; the
text moved behind --debug.

Both are covered by mutation-checked tests, and the README gains a Security
section covering the unauthenticated trust boundary, the 0.0.0.0 bind that
also exposes VPN interfaces, and running the daemon as a non-admin user.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-29 03:44:25 +01:00
co-authored by Claude Opus 5
parent 9bf887e780
commit 5a6fc62106
4 changed files with 170 additions and 9 deletions
+41
View File
@@ -302,6 +302,47 @@ tail -f /tmp/local.wyoming-parakeet.stderr
Each request logs audio duration, inference time and the transcript. Each request logs audio duration, inference time and the transcript.
## Security
**The Wyoming protocol has no authentication or transport encryption.** Any
client that can reach the port can submit audio, and Home Assistant trusts
whatever this service returns — a transcript becomes an intent, and an intent
unlocks doors. Treat the port as a trust boundary and keep it on a network you
control.
Two things this server does about that:
- **Buffered audio is capped** (`--max-audio-seconds`, default 120). Audio
accumulates until `AudioStop`, and a client that never sends one — malicious
or just stuck — otherwise grows the buffer indefinitely. Measured at ~11 MB/s
over loopback, enough to exhaust 32 GB in under an hour from one connection.
Over the cap, further audio is dropped with a single warning and whatever was
captured is still transcribed.
- **Transcripts are not logged at INFO.** The log records duration, latency and
character count; the text itself is behind `--debug`. Log files are
long-lived, and on macOS `/tmp` they are world-readable by default — meaning
every voice command would otherwise sit in plaintext readable by any local
account.
Worth doing yourself, depending on your threat model:
- **Bind to one interface.** The daemon listens on `0.0.0.0`, so it is exposed
on every network the host is attached to — including VPN interfaces like
Tailscale, which is easy to overlook. Set the URI to a specific address
(`--uri tcp://192.168.1.10:7892`) or firewall the port to your Home
Assistant host.
- **Do not run it as an admin account.** The installer defaults `--user` to
whoever runs it. On a typical macOS setup that account is in `admin`, and if
`%admin` has a `NOPASSWD` sudo rule then a compromise of this service is a
direct path to root. A dedicated non-admin service account costs nothing.
- **Pin your dependencies** if you care about supply chain. `requirements.txt`
is deliberately loose so `install.sh` picks up fixes; pin exact versions (and
ideally hashes) if you would rather audit upgrades. Model weights are
`safetensors`, so loading them does not execute code, but the initial
download from HuggingFace is trust-on-first-use — pin a revision if that
matters to you. `HF_HUB_OFFLINE=1` in the daemon means it never re-fetches
after that point.
## License ## License
MIT MIT
+88 -4
View File
@@ -1,4 +1,5 @@
"""Tests for the Wyoming event handling.""" """Tests for the Wyoming event handling."""
from argparse import Namespace
from unittest.mock import AsyncMock, MagicMock from unittest.mock import AsyncMock, MagicMock
import pytest import pytest
@@ -7,7 +8,12 @@ from wyoming.audio import AudioChunk, AudioStop
from wyoming.info import AsrModel, AsrProgram, Attribution, Describe, Info from wyoming.info import AsrModel, AsrProgram, Attribution, Describe, Info
from wyoming_parakeet.engine import SAMPLE_RATE from wyoming_parakeet.engine import SAMPLE_RATE
from wyoming_parakeet.handler import ParakeetEventHandler from wyoming_parakeet.handler import (
DEFAULT_MAX_AUDIO_SECONDS,
ParakeetEventHandler,
)
ARGS = Namespace(max_audio_seconds=DEFAULT_MAX_AUDIO_SECONDS)
INFO = Info( INFO = Info(
asr=[ asr=[
@@ -36,7 +42,7 @@ INFO = Info(
def handler(): def handler():
engine = MagicMock() engine = MagicMock()
engine.transcribe = AsyncMock(return_value="turn off the kitchen lights") engine.transcribe = AsyncMock(return_value="turn off the kitchen lights")
h = ParakeetEventHandler(INFO, MagicMock(), engine, MagicMock(), MagicMock()) h = ParakeetEventHandler(INFO, ARGS, engine, MagicMock(), MagicMock())
h.write_event = AsyncMock() h.write_event = AsyncMock()
return h return h
@@ -172,8 +178,8 @@ async def test_concurrent_handlers_do_not_share_audio():
seen = [] seen = []
engine.transcribe.side_effect = lambda pcm: seen.append(pcm) or "" engine.transcribe.side_effect = lambda pcm: seen.append(pcm) or ""
a = ParakeetEventHandler(INFO, MagicMock(), engine, MagicMock(), MagicMock()) a = ParakeetEventHandler(INFO, ARGS, engine, MagicMock(), MagicMock())
b = ParakeetEventHandler(INFO, MagicMock(), engine, MagicMock(), MagicMock()) b = ParakeetEventHandler(INFO, ARGS, engine, MagicMock(), MagicMock())
a.write_event = AsyncMock() a.write_event = AsyncMock()
b.write_event = AsyncMock() b.write_event = AsyncMock()
@@ -224,3 +230,81 @@ async def test_8bit_input_is_widened(handler):
await handler.handle_event(AudioStop().event()) await handler.handle_event(AudioStop().event())
assert len(handler.engine.transcribe.await_args[0][0]) == SAMPLE_RATE * 2 assert len(handler.engine.transcribe.await_args[0][0]) == SAMPLE_RATE * 2
# --- resource limits --------------------------------------------------------
# Wyoming is unauthenticated, so buffered audio is attacker-controlled.
def capped_handler(seconds):
engine = MagicMock()
engine.transcribe = AsyncMock(return_value="ok")
h = ParakeetEventHandler(
INFO, Namespace(max_audio_seconds=seconds), engine, MagicMock(), MagicMock()
)
h.write_event = AsyncMock()
return h
async def test_audio_buffer_is_capped():
handler = capped_handler(1.0)
for _ in range(10):
await handler.handle_event(chunk(SAMPLE_RATE)) # 10s into a 1s cap
assert len(handler.audio) == SAMPLE_RATE * 2
async def test_truncated_audio_is_still_transcribed():
"""Cap the memory, don't drop the user's command on the floor."""
handler = capped_handler(1.0)
for _ in range(5):
await handler.handle_event(chunk(SAMPLE_RATE))
await handler.handle_event(AudioStop().event())
handler.engine.transcribe.assert_awaited_once()
assert len(handler.engine.transcribe.await_args[0][0]) == SAMPLE_RATE * 2
assert Transcript.from_event(handler.write_event.await_args[0][0]).text == "ok"
async def test_truncation_warns_once(caplog):
handler = capped_handler(1.0)
with caplog.at_level("WARNING"):
for _ in range(6):
await handler.handle_event(chunk(SAMPLE_RATE))
assert caplog.text.count("Audio exceeded") == 1
async def test_cap_resets_between_utterances():
handler = capped_handler(1.0)
for _ in range(3):
await handler.handle_event(chunk(SAMPLE_RATE))
await handler.handle_event(AudioStop().event())
assert handler.truncated is False
await handler.handle_event(chunk(SAMPLE_RATE // 2))
assert len(handler.audio) == SAMPLE_RATE # accepted again, not still capped
async def test_audio_under_the_cap_is_untouched():
handler = capped_handler(DEFAULT_MAX_AUDIO_SECONDS)
await handler.handle_event(chunk(SAMPLE_RATE * 3))
assert len(handler.audio) == SAMPLE_RATE * 3 * 2
# --- transcript privacy -----------------------------------------------------
# Logs are long-lived and readable by other local accounts.
async def test_transcript_text_is_not_logged_at_info(handler, caplog):
handler.engine.transcribe = AsyncMock(return_value="unlock the front door")
with caplog.at_level("INFO"):
await handler.handle_event(chunk(SAMPLE_RATE))
await handler.handle_event(AudioStop().event())
assert "unlock the front door" not in caplog.text
async def test_transcript_text_is_available_at_debug(handler, caplog):
handler.engine.transcribe = AsyncMock(return_value="unlock the front door")
with caplog.at_level("DEBUG"):
await handler.handle_event(chunk(SAMPLE_RATE))
await handler.handle_event(AudioStop().event())
assert "unlock the front door" in caplog.text
+12 -2
View File
@@ -8,7 +8,7 @@ from wyoming.info import AsrModel, AsrProgram, Attribution, Info
from wyoming.server import AsyncServer from wyoming.server import AsyncServer
from .engine import ParakeetEngine from .engine import ParakeetEngine
from .handler import ParakeetEventHandler from .handler import DEFAULT_MAX_AUDIO_SECONDS, ParakeetEventHandler
_LOGGER = logging.getLogger(__name__) _LOGGER = logging.getLogger(__name__)
__version__ = "1.0.0" __version__ = "1.0.0"
@@ -25,7 +25,17 @@ def build_parser() -> argparse.ArgumentParser:
parser.add_argument("--uri", required=True, help="unix:// or tcp://") parser.add_argument("--uri", required=True, help="unix:// or tcp://")
parser.add_argument("--model", default=DEFAULT_MODEL, help="HuggingFace model id") parser.add_argument("--model", default=DEFAULT_MODEL, help="HuggingFace model id")
parser.add_argument("--language", default="en", help="Language code reported to HA") parser.add_argument("--language", default="en", help="Language code reported to HA")
parser.add_argument("--debug", action="store_true", help="Log DEBUG messages") parser.add_argument(
"--max-audio-seconds",
type=float,
default=DEFAULT_MAX_AUDIO_SECONDS,
help="Cap on buffered audio per utterance (default: %(default)s)",
)
parser.add_argument(
"--debug",
action="store_true",
help="Log DEBUG messages, including transcript text",
)
parser.add_argument("--log-format", default=logging.BASIC_FORMAT) parser.add_argument("--log-format", default=logging.BASIC_FORMAT)
parser.add_argument("--version", action="version", version=__version__) parser.add_argument("--version", action="version", version=__version__)
return parser return parser
+29 -3
View File
@@ -12,6 +12,13 @@ from .engine import SAMPLE_RATE
_LOGGER = logging.getLogger(__name__) _LOGGER = logging.getLogger(__name__)
# Wyoming has no authentication, so any client that can reach the port can
# stream audio. Without a cap, self.audio grows until AudioStop -- measured at
# ~11 MB/s over loopback, which exhausts 32 GB in under an hour from a single
# connection. A stuck satellite that never sends AudioStop does the same thing
# by accident. No real voice command approaches this bound.
DEFAULT_MAX_AUDIO_SECONDS = 120
class ParakeetEventHandler(AsyncEventHandler): class ParakeetEventHandler(AsyncEventHandler):
def __init__(self, wyoming_info: Info, cli_args, engine, *args, **kwargs): def __init__(self, wyoming_info: Info, cli_args, engine, *args, **kwargs):
@@ -20,13 +27,27 @@ class ParakeetEventHandler(AsyncEventHandler):
self.wyoming_info_event = wyoming_info.event() self.wyoming_info_event = wyoming_info.event()
self.engine = engine self.engine = engine
self.audio = bytes() self.audio = bytes()
self.truncated = False
self.converter = AudioChunkConverter(rate=SAMPLE_RATE, width=2, channels=1) self.converter = AudioChunkConverter(rate=SAMPLE_RATE, width=2, channels=1)
max_seconds = getattr(cli_args, "max_audio_seconds", DEFAULT_MAX_AUDIO_SECONDS)
self.max_bytes = int(max_seconds * SAMPLE_RATE * 2)
def _append(self, chunk: bytes) -> None:
room = self.max_bytes - len(self.audio)
if room > 0:
self.audio += chunk[:room]
if len(self.audio) >= self.max_bytes and not self.truncated:
self.truncated = True
_LOGGER.warning(
"Audio exceeded %.0fs; ignoring the rest of this utterance",
self.max_bytes / (SAMPLE_RATE * 2),
)
async def handle_event(self, event: Event) -> bool: async def handle_event(self, event: Event) -> bool:
if AudioChunk.is_type(event.type): if AudioChunk.is_type(event.type):
if not self.audio: if not self.audio:
_LOGGER.debug("Receiving audio") _LOGGER.debug("Receiving audio")
self.audio += self.converter.convert(AudioChunk.from_event(event)).audio self._append(self.converter.convert(AudioChunk.from_event(event)).audio)
return True return True
if AudioStop.is_type(event.type): if AudioStop.is_type(event.type):
@@ -34,12 +55,16 @@ class ParakeetEventHandler(AsyncEventHandler):
started = time.monotonic() started = time.monotonic()
try: try:
text = await self.engine.transcribe(self.audio) text = await self.engine.transcribe(self.audio)
# The transcript is everything the user said, and the log file
# is long-lived and readable by other local accounts. Keep the
# operational signal at INFO and the content behind --debug.
_LOGGER.info( _LOGGER.info(
"%.2fs audio -> %.0fms :: %r", "%.2fs audio -> %.0fms, %d chars",
duration, duration,
(time.monotonic() - started) * 1000, (time.monotonic() - started) * 1000,
text, len(text),
) )
_LOGGER.debug("Transcript: %r", text)
except Exception: except Exception:
# wyoming's run loop has no except clause, so letting this # wyoming's run loop has no except clause, so letting this
# propagate closes the connection without ever sending a # propagate closes the connection without ever sending a
@@ -49,6 +74,7 @@ class ParakeetEventHandler(AsyncEventHandler):
text = "" text = ""
finally: finally:
self.audio = bytes() self.audio = bytes()
self.truncated = False
await self.write_event(Transcript(text=text).event()) await self.write_event(Transcript(text=text).event())
return False return False