Read the stations that never say a word
Most of what identifies itself on the air identifies itself in Morse. A repeater, a beacon, an unattended transmitter: four to six characters, over in a second or two, and no speech anywhere in the capture. Every one of those was being thrown away, in three separate places. The callsign book and the map were built inside the transcription branch, on the reasoning that callsigns come out of transcripts. They also come out of Morse and out of APRS headers, neither of which involves a speech recogniser -- so a receiver with none installed found none of them, and a CW ident reached the sidecar and stopped there. Both are now built whenever classification is on, and all three sources go through one place. The CW decoder only ran where the classifier had already said cw, ook or carrier. A two-second ident is a fraction of a capture named after whatever filled the rest of it. Every capture is offered to it now, once it has finished; a decode does not relabel a capture that plainly holds speech. And the decoder's own gates were written for a paragraph. Three characters, eight elements, and any repeated character refused -- which read VVV, DE, AR and K correctly and then discarded them. Short is the normal case now, on a second bar: perfect timing, nothing undecoded, and the keyed tone at least 20 dB over its band. That last is not decoration. With four elements the dot length is fitted to those very elements, so noise lands on the grid as neatly as keying does; a third of a second of white noise decodes as a perfectly timed V. Over 200 noise blocks the loudest bin never rose 13 dB above the median while keying at 3 dB SNR sits above 40. One keyed element is still refused: a single pulse is an E or a T whether a person sent it or the squelch opened on a click. 288 non-Morse cases, no false positives. Feeding that text to a callsign lookup made truncation matter. A capture opens when the squelch does, halfway through an element as often as not, and half a character is not a smaller reading -- a K missing its first dash is an A. So the sliced character is dropped, and so is the rest of its word, because what is left can read as a whole one: K1AA caught halfway through is K1A, which is somebody else. Across 1805 truncated captures that is 107 invented callsigns down to none, with 550 correct ones still found. Phonetics, which is the other half of the ask. A recogniser has never heard of the alphabet -- it writes what the words sounded like: Whiskey-One-Alpha-Whiskey hyphenated WhiskeyOneAlphaWhiskey run together Whiskey1AlphaWhiskey and half in digits wiskey one alfa whisky spelled the way it sounded whiskey one alpha, uh, whiskey with the hesitation written down All read back to W1AW now. A word is only taken apart when it is phonetic all the way through, which is what keeps it off "kilometre" and "victorious". Two bugs found on the way, both of which invented a callsign: - Nothing is joined across a slash any more. The beacon W1AW/B came back as W1AWB, which belongs to nobody, and W1AW-4 came back as nothing at all. - Nor across a gap the sender chose. A transcript's spacing is the recogniser's guess and may be closed up; a word gap in Morse is seven dot units, so "KU0W K" is a station signing off, not a longer callsign. Also here: - saunterbrowse gives Morse a panel of its own, with the licence under it, searchable with / and readable with t. - --simulate no longer looks anything up or writes a map. The demo band is invented but W1AW is the ARRL's own station, and it would have been pinned to the same map a real scan writes. - classify._psk_order took the logarithm of zero on a silent block. - The demo band has a repeater ident in it, and its Morse no longer runs one repeat into the next. - conftest refuses a real licence lookup from any test. 1161 tests, up from 1014. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PsWPTweCT6pwxKngvVxcg
This commit is contained in:
parent
f9f0d94000
commit
b4718aa425
21 changed files with 1434 additions and 129 deletions
76
README.md
76
README.md
|
|
@ -773,7 +773,51 @@ reliably:
|
|||
```
|
||||
|
||||
The decoder runs its own CW detector over the captured IQ, so Morse is found
|
||||
even when the recording itself was made in FM or SSB.
|
||||
even when the recording itself was made in FM or SSB — and it is run over
|
||||
**every** capture once it has finished, whatever the classifier called it.
|
||||
|
||||
#### Short bursts, which is most of it
|
||||
|
||||
Most of the Morse on the air is not a conversation. It is a repeater, a beacon
|
||||
or an unattended transmitter saying who it is and stopping — four to six
|
||||
characters, over in a second or two:
|
||||
|
||||
```
|
||||
147.06 MHz 5.3s SNR 31.2 dB CW / Morse at 20 WPM CW "DE K1AA"
|
||||
K1AA Newington Radio Club — Newington, CT · FN31pr
|
||||
```
|
||||
|
||||
That burst is a fraction of a capture the classifier named after whatever
|
||||
filled the rest of it, so waiting for the label to say "CW" missed it. Short
|
||||
is now the normal case rather than the awkward one, and a decode of two or
|
||||
three characters gets in on its timing alone:
|
||||
|
||||
- every element within a third of a unit of one or three
|
||||
- every character resolving to something in the table
|
||||
- **and the keyed tone at least 20 dB above the rest of its band**
|
||||
|
||||
The last one is what separates an ident from a blip, and it is not
|
||||
decoration. With four elements the dot length is fitted to those very
|
||||
elements, so they land on the grid whatever produced them — a third of a
|
||||
second of white noise decodes as a perfectly timed `V`. Measured over 200
|
||||
noise blocks the loudest bin never rose 13 dB above the median of its band,
|
||||
while keying at 3 dB SNR sits above 40, so 20 dB has room on both sides. One
|
||||
keyed element is refused outright: a single pulse is an `E` or a `T` whether a
|
||||
person sent it or the squelch opened on a click.
|
||||
|
||||
#### What the capture window cut off
|
||||
|
||||
A capture opens when the squelch does, which is in the middle of an element as
|
||||
often as not. Half a character is not a smaller reading of what was sent — it
|
||||
is a different one. A `K` missing its first dash is an `A`; a `W` missing its
|
||||
first dot is an `M`.
|
||||
|
||||
So the character at a sliced end is dropped, and so is the rest of the word it
|
||||
was in, because what is left of that word can read as a whole one: **`K1AA`
|
||||
caught halfway through is `K1A`, which belongs to somebody else.** The full
|
||||
text is still shown; it is the *identification* that is held to the stricter
|
||||
standard. Over 1805 truncated captures of four different messages, that turns
|
||||
107 invented callsigns into none, while still recovering 550 correct ones.
|
||||
|
||||
### Single sideband
|
||||
|
||||
|
|
@ -1209,7 +1253,35 @@ Note what the recogniser actually wrote: **"KU 0W"**, with a space. Speech
|
|||
recognisers are poor at callsigns — they are not words, they are said one
|
||||
character at a time — so a callsign arrives broken wherever the speaker
|
||||
paused, and an operator who spells it out gets *"kilo uniform zero whiskey"*
|
||||
written down verbatim. All three forms read back to `KU0W`.
|
||||
written down verbatim.
|
||||
|
||||
A recogniser has never heard of the phonetic alphabet, so it writes what the
|
||||
words sounded like and does whatever it likes with the spacing. All of these
|
||||
are one callsign, and all of them read back correctly:
|
||||
|
||||
| What the recogniser wrote | Why |
|
||||
|---|---|
|
||||
| `KU 0W`, `K7 RA` | broken where the speaker paused |
|
||||
| `kilo uniform zero whiskey` | spelled out, one word per character |
|
||||
| `Whiskey-One-Alpha-Whiskey` | spelled out and hyphenated |
|
||||
| `WhiskeyOneAlphaWhiskey`, `Whiskey1AlphaWhiskey` | run together |
|
||||
| `wiskey one alfa whisky` | spelled the way it sounded |
|
||||
| `whiskey one alpha, uh, whiskey` | said with a hesitation in the middle |
|
||||
| `W1AW-4`, `W1AW/B`, `DL/W1AW` | a suffix, which is not part of the callsign |
|
||||
|
||||
A word is only taken apart when it is phonetic *all the way through*, which is
|
||||
what keeps this away from English: "kilometre" begins with a phonetic word and
|
||||
"victorious" contains one, and neither can be consumed to the end.
|
||||
|
||||
Callsigns arrive from three directions and all three end up in the same list
|
||||
and on the same map: **spoken and transcribed, sent in Morse, or carried in
|
||||
the header of an APRS packet.** Neither of the last two involves a speech
|
||||
recogniser, so a machine with none installed still builds a map.
|
||||
|
||||
Where the gaps came from decides whether they can be closed. A transcript's
|
||||
spacing is the recogniser's guess, so `KU 0W` may be joined; a word gap in
|
||||
Morse is seven dot units the sender chose, so `KU0W K` is a station signing
|
||||
off, not a callsign one letter longer.
|
||||
|
||||
The other half of the problem is not inventing them. A browser that reports
|
||||
callsigns nobody said is worse than one that reports none, so a run of words
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue