<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Posts on Douglas Santos</title><link>https://blog.batcode.io/en/posts/</link><description>Recent content in Posts on Douglas Santos</description><generator>Hugo -- 0.166.0</generator><language>en-US</language><lastBuildDate>Wed, 09 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://blog.batcode.io/en/posts/index.xml" rel="self" type="application/rss+xml"/><item><title>Ghost Bluetooth on Linux: how btusb wedges the MediaTek MT6639 trying to load firmware that isn't there.</title><link>https://blog.batcode.io/en/posts/mt6639-bluetooth-reset-loop/</link><pubDate>Wed, 09 Sep 2026 00:00:00 +0000</pubDate><guid>https://blog.batcode.io/en/posts/mt6639-bluetooth-reset-loop/</guid><description>Bluetooth never worked on this board and I was sure the missing firmware was the problem. It was — but what actually broke the hardware was the driver trying to fix it forever. Here is the whole investigation, from a false negative in dmesg to a 174 millisecond race at boot.</description><content:encoded><![CDATA[<p>I bought an ASUS ProArt X870E-Creator WiFi and Wi-Fi worked on the first boot. Bluetooth did not. No adapter ever showed up, and worse: the device was not in <code>lsusb</code> at all. Not failing, not erroring — absent.</p>
<p>My first guess was the obvious one, and it was half right: the Bluetooth firmware for this chip does not ship in <code>linux-firmware</code>. What I did not expect is that the missing firmware is not what breaks the hardware. What breaks it is what the driver does about it.</p>
<p>In this article I will walk through the whole investigation, including the two moments where I confidently reached the wrong conclusion. I think those parts are more useful than the fix.</p>
<table>
	<thead>
			<tr>
					<th></th>
					<th></th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Board</td>
					<td>ASUS ProArt X870E-Creator WiFi rev 2</td>
			</tr>
			<tr>
					<td>BIOS</td>
					<td>2402</td>
			</tr>
			<tr>
					<td>OS</td>
					<td>Bazzite (Fedora 44 Atomic)</td>
			</tr>
			<tr>
					<td>Kernel</td>
					<td>7.2.3-ogc3.1.fc44.x86_64</td>
			</tr>
			<tr>
					<td>Bluetooth</td>
					<td>MediaTek MT6639, USB <code>0489:e13a</code></td>
			</tr>
			<tr>
					<td>Wi-Fi</td>
					<td>MediaTek MT7927, PCIe <code>14c3:7927</code></td>
			</tr>
	</tbody>
</table>
<h2 id="the-symptom-an-adapter-that-does-not-exist">The symptom: an adapter that does not exist</h2>
<p>No adapter. <code>bluetoothctl show</code> hangs without printing anything, because bluetoothd sits waiting on a controller that is never going to appear. And <code>dmesg</code> had absolutely no messages from the Bluetooth subsystem — no success, no firmware failure, no USB error. Complete silence.</p>
<p>Complete silence is a strange result. If the firmware were missing, I should see the driver complaining. If the device were faulty, I should see USB complaining. Seeing nothing means the kernel never even tried.</p>
<h2 id="the-false-negative-that-cost-me-a-whole-round">The false negative that cost me a whole round</h2>
<p>Here is the first mistake, and it is embarrassing in how simple it is.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">dmesg <span class="p">|</span> grep -iE <span class="s1">&#39;btusb|bluetooth: hci&#39;</span>
</span></span></code></pre></div><p>This came back empty. I read &ldquo;empty&rdquo; as &ldquo;there are no Bluetooth messages&rdquo;. Wrong. What was actually happening is that <code>kernel.dmesg_restrict=1</code> makes <code>dmesg</code> return <strong>zero lines of any kind</strong> without root. There were no messages at all to filter — not about Bluetooth, not about USB, not about anything.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">$ dmesg <span class="p">|</span> wc -l
</span></span><span class="line"><span class="cl"><span class="m">0</span>
</span></span><span class="line"><span class="cl">$ sysctl -n kernel.dmesg_restrict
</span></span><span class="line"><span class="cl"><span class="m">1</span>
</span></span></code></pre></div><p>An empty <code>grep</code> over empty input looks exactly like an empty <code>grep</code> over a thousand lines. The practical lesson: use <code>journalctl -k</code>, which works without sudo and gives you the real log.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">$ journalctl -k -b <span class="m">0</span> --no-pager <span class="p">|</span> wc -l
</span></span><span class="line"><span class="cl"><span class="m">1893</span>
</span></span></code></pre></div><p>One thousand eight hundred and ninety three lines I had been ignoring. And everything was in there.</p>
<h2 id="the-device-did-not-disappear--it-wedged">The device did not disappear — it wedged</h2>
<p>With the real log in hand the picture changed completely. The device <strong>is</strong> there, on port <code>1-6</code>, right next to the board&rsquo;s LED controller. It is detected electrically. It just does not answer anything.</p>
<pre tabindex="0"><code>usb 1-6: new high-speed USB device number 4 using xhci_hcd
usb 1-6: device descriptor read/64, error -110
usb 1-6: device descriptor read/64, error -110
usb 1-6: new high-speed USB device number 5 using xhci_hcd
usb 1-6: device descriptor read/64, error -110
usb 1-6: device descriptor read/64, error -110
usb usb1-port6: attempt power cycle
usb 1-6: new high-speed USB device number 6 using xhci_hcd
usb 1-6: Device not responding to setup address.
usb 1-6: device not accepting address 6, error -71
usb 1-6: new high-speed USB device number 7 using xhci_hcd
usb 1-6: Device not responding to setup address.
usb 1-6: device not accepting address 7, error -71
usb usb1-port6: unable to enumerate USB device
</code></pre><p>Translating: <code>-110</code> is a timeout, <code>-71</code> is a protocol error. The hub sees something plugged in, tries to read the device descriptor, gets no answer, cuts and restores power to the port, tries again, gives up. Four attempts, sixty three seconds, nothing.</p>
<p>This also explains why <code>btusb</code> was not loaded, and why that was <strong>not</strong> the problem. udev loads the module when a matching modalias shows up. Since nothing enumerated, there is no modalias, so the module does not load. Loading it by hand works with no errors and creates no HCI device at all:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">$ sudo modprobe btusb <span class="o">&amp;&amp;</span> ls /sys/class/bluetooth/
</span></span><span class="line"><span class="cl"><span class="c1"># (empty)</span>
</span></span></code></pre></div><p>The software stack is spotless. It is the hardware that never shows up.</p>
<h2 id="the-cause-btusb-retries-forever">The cause: btusb retries forever</h2>
<p>Bazzite keeps previous boots in the journal, and that is where it gets interesting. I swept every recorded boot looking for the device, and in two of them it <strong>worked</strong> — it enumerated normally. So I went to look at what had happened:</p>
<pre tabindex="0"><code>[    3.068951] usb 1-6: New USB device found, idVendor=0489, idProduct=e13a
[    3.069092] usb 1-6: Product: Wireless_Device
[    8.221018] usbcore: registered new interface driver btusb
[    8.233139] Bluetooth: hci0: Failed to load firmware file (-2)
[    8.233145] Bluetooth: hci0: Failed to set up firmware (-2)
[    8.564037] usb 1-6: reset high-speed USB device number 4 using xhci_hcd
[    8.817354] Bluetooth: hci0: Failed to load firmware file (-2)
[    9.147127] usb 1-6: reset high-speed USB device number 4 using xhci_hcd
[    9.402354] Bluetooth: hci0: Failed to load firmware file (-2)
[    9.727020] usb 1-6: reset high-speed USB device number 4 using xhci_hcd
</code></pre><p>See the pattern? The firmware fails with <code>-2</code> (ENOENT, file not found), and <code>btusb</code> <strong>resets the device over USB and tries again</strong>. Then it fails again, resets again, tries again. Every 0.58 seconds. No backoff, no retry limit, never gives up.</p>
<p>On that boot this ran for 13 minutes until I shut the machine down. That was 1335 resets.</p>
<p>And that is what wedges the chip. The full chain goes like this:</p>
<ol>
<li>Clean cold start: the controller enumerates normally, about 3 seconds into boot.</li>
<li><code>btusb</code> binds at around 8 seconds, creates <code>hci0</code>, requests the firmware.</li>
<li>The file is not there. It returns <code>-2</code>.</li>
<li><strong><code>btusb</code> resets the device and tries again. And again. Indefinitely.</strong></li>
<li>After a few hundred resets, the controller firmware wedges.</li>
<li>From then on the port detects the device but it no longer answers anything.</li>
<li>This <strong>survives reboots</strong>, because the standby power rail keeps the chip alive.</li>
</ol>
<p>Point 7 is what makes it cruel. You reboot, no change. You reinstall the OS, no change. The chip stays wedged because it was never actually powered down.</p>
<h2 id="the-numbers-line-up">The numbers line up</h2>
<p>What convinced me this was really the cause, and not a coincidence, was counting. If the reset loop is a consequence of the firmware failure, the two numbers have to be identical:</p>
<table>
	<thead>
			<tr>
					<th>Boot</th>
					<th style="text-align: right">Resets of port 1-6</th>
					<th style="text-align: right">Firmware failures</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>-8</td>
					<td style="text-align: right">398</td>
					<td style="text-align: right">398</td>
			</tr>
			<tr>
					<td>-1</td>
					<td style="text-align: right">1335</td>
					<td style="text-align: right">1335</td>
			</tr>
	</tbody>
</table>
<p>They match exactly. One reset per failure, in both boots.</p>
<p>And the recurrence pattern across boots tells the rest of the story:</p>
<table>
	<thead>
			<tr>
					<th>Boot</th>
					<th>Port 1-6</th>
					<th>Reading</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>-8</td>
					<td>enumerated</td>
					<td>clean start, 398 resets, wedges</td>
			</tr>
			<tr>
					<td>-7 to -3</td>
					<td>no events</td>
					<td><strong>five dead boots</strong>, caused by -8</td>
			</tr>
			<tr>
					<td>-2</td>
					<td>failed to enumerate</td>
					<td>wedged</td>
			</tr>
			<tr>
					<td>-1</td>
					<td>enumerated</td>
					<td>freed by a power cut, 1335 resets, wedges again</td>
			</tr>
			<tr>
					<td>0</td>
					<td>failed to enumerate</td>
					<td>wedged</td>
			</tr>
	</tbody>
</table>
<p>Two complete cycles. Every time I freed the chip by cutting power, it came back, entered the loop, and wedged again. I was reproducing the problem without knowing it.</p>
<p>To free it, you have to actually cut power: enable <strong>ErP in S4+S5</strong> in the BIOS and use <code>poweroff</code> (not <code>reboot</code>, which never cuts standby), or unplug the machine for about 10 seconds.</p>
<h2 id="by-the-way-the-slow-boot-was-the-same-bug">By the way, the slow boot was the same bug</h2>
<p>I had a second complaint I thought was unrelated: boot was taking almost two minutes. It was related.</p>
<p><code>systemd-udev-settle.service</code> sits at the <strong>root of the boot&rsquo;s critical chain</strong> — everything waits on it. And the enumeration retries on port 1-6 keep udev busy:</p>
<table>
	<thead>
			<tr>
					<th>State of port 1-6</th>
					<th>Boots</th>
					<th style="text-align: right">udev-settle</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>device absent</td>
					<td>-7 to -3</td>
					<td style="text-align: right">8.6 - 9.7 s</td>
			</tr>
			<tr>
					<td>device wedged</td>
					<td>-2, 0</td>
					<td style="text-align: right"><strong>67.8 s</strong></td>
			</tr>
	</tbody>
</table>
<p>About 59 seconds of penalty, matching the 63 second window of the retries. Once the controller came up properly, boot went from <strong>1 min 52 s to 41 s</strong>. Two symptoms, one bug.</p>
<h2 id="where-to-put-firmware-when-usr-is-read-only">Where to put firmware when /usr is read-only</h2>
<p>Bazzite is an immutable system: <code>/usr</code> is read-only, so you cannot just drop the file into <code>/usr/lib/firmware</code>. The standard route is to use <code>/var/lib/firmware</code> and point the kernel at it with a boot parameter:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">sudo install -Dm644 BT_RAM_CODE_MT6639_2_1_hdr.bin <span class="se">\
</span></span></span><span class="line"><span class="cl">  /var/lib/firmware/mediatek/mt7927/BT_RAM_CODE_MT6639_2_1_hdr.bin
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">sudo rpm-ostree kargs --append<span class="o">=</span>firmware_class.path<span class="o">=</span>/var/lib/firmware
</span></span></code></pre></div><p>I did that, rebooted, and Bluetooth worked. End of article, right?</p>
<p>No. I went to check the log and the firmware had <strong>failed with <code>-2</code> again</strong> — and the device came up anyway. That made no sense at all, and the explanation is the most interesting part of the whole thing.</p>
<h2 id="the-race-i-won-by-174-milliseconds">The race I won by 174 milliseconds</h2>
<p>On ostree systems, <code>/var</code> is a separate subvolume mounted by a systemd unit <strong>after</strong> switch-root. And <code>btusb</code> probes right inside that window.</p>
<table>
	<thead>
			<tr>
					<th style="text-align: right">Time</th>
					<th>Event</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td style="text-align: right">7.294 s</td>
					<td>switch-root</td>
			</tr>
			<tr>
					<td style="text-align: right">8.677 s</td>
					<td><code>btusb</code> registers the interface driver</td>
			</tr>
			<tr>
					<td style="text-align: right"><strong>8.688 s</strong></td>
					<td><strong>first firmware attempt fails, <code>-2</code></strong></td>
			</tr>
			<tr>
					<td style="text-align: right">9.021 s</td>
					<td><code>btusb</code> resets the device and reschedules</td>
			</tr>
			<tr>
					<td style="text-align: right"><strong>9.235 s</strong></td>
					<td><strong><code>var.mount</code> completes</strong></td>
			</tr>
			<tr>
					<td style="text-align: right">9.409 s</td>
					<td>the retry <strong>finds</strong> the firmware</td>
			</tr>
			<tr>
					<td style="text-align: right">28.758 s</td>
					<td><code>Device setup in 19036182 usecs</code></td>
			</tr>
			<tr>
					<td style="text-align: right">28.930 s</td>
					<td><code>AOSP extensions version v1.00</code></td>
			</tr>
	</tbody>
</table>
<p>So it worked for the wrong reason. The first iteration of the reset loop — the same loop that wedges the chip — is exactly what saved it, because <code>/var</code> mounted in the middle of it. The margin was 174 milliseconds.</p>
<p>That is reproducible luck, not a guarantee. If the retry landed 200 ms earlier, the loop would start and the chip would wedge. And there is a cruel detail on top: there is a non-empty <code>/var</code> stub underneath the mount point in the deployment, so the lookup fails <strong>silently</strong> instead of reporting a missing directory.</p>
<h2 id="the-real-fix">The real fix</h2>
<p>The solution is to put the firmware somewhere already readable at switch-root. <code>/etc</code> qualifies: it is a bind mount of the deployment&rsquo;s own subvolume, set up by <code>ostree-prepare-root</code> while still inside the initrd. It is available at 7.294 s, more than a second before <code>btusb</code> probes.</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">sudo install -Dm644 BT_RAM_CODE_MT6639_2_1_hdr.bin <span class="se">\
</span></span></span><span class="line"><span class="cl">  /etc/firmware/mediatek/mt7927/BT_RAM_CODE_MT6639_2_1_hdr.bin
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">sudo rpm-ostree kargs <span class="se">\
</span></span></span><span class="line"><span class="cl">  --replace<span class="o">=</span>firmware_class.path<span class="o">=</span>/var/lib/firmware<span class="o">=</span>/etc/firmware
</span></span></code></pre></div><p>A nice detail: SELinux labels the file <code>cpucontrol_conf_t</code>, because <code>/etc/firmware</code> is already a path the policy knows about (used for CPU microcode). And it blocks nothing — the loaded policy does not even define the <code>firmware_load</code> permission.</p>
<p>To validate after the reboot, both numbers have to come out zero:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">journalctl -k -b <span class="m">0</span> <span class="p">|</span> grep -c <span class="s1">&#39;Failed to load firmware file&#39;</span>
</span></span><span class="line"><span class="cl">journalctl -k -b <span class="m">0</span> <span class="p">|</span> grep -c <span class="s1">&#39;reset high-speed USB device&#39;</span>
</span></span></code></pre></div><h2 id="the-staged-etc-trap--i-almost-gave-up-here">The staged /etc trap — I almost gave up here</h2>
<p>This was the second moment where I confidently reached the wrong conclusion, and it was worth a scare.</p>
<p>After <code>rpm-ostree kargs</code>, I went to check the new deployment before rebooting. <code>/etc/firmware</code> <strong>was not there</strong>. And the deployment is read-only, so I could not even put it there by hand.</p>
<p>This looked fatal. With the boot parameter now pointing only at <code>/etc/firmware</code>, a missing file means the first attempt <strong>and the retry</strong> both fail, since <code>/var</code> is no longer in the search path. In other words: worse than not touching anything.</p>
<p>Before reverting, I decided to compare the two <code>/etc</code> trees. The new deployment&rsquo;s was missing 23 entries the running one had:</p>
<pre tabindex="0"><code>bazzite    cardwire   cni        crypttab   firmware
fstab      group-     gshadow-   hostname   iwd
locale.conf localtime passwd-    sddm.conf.d shadow-
subgid-    subuid-    vconsole.conf
</code></pre><p>Look at <code>fstab</code> sitting there in the middle.</p>
<p>That is what settles it. If finalization did not merge <code>/etc</code>, <strong>no ostree system could boot after an upgrade</strong>, because it would have no <code>fstab</code>. Therefore, a staged deployment&rsquo;s <code>/etc</code> is the pristine one from the new commit, and the three-way merge is deferred to <code>ostree admin finalize-staged</code>, which runs at shutdown.</p>
<p>In other words: inspecting a staged <code>/etc</code> will <strong>always</strong> look like your local changes are gone. That is expected, not a defect. My file was going to be carried over along with <code>fstab</code> and <code>hostname</code>, and it was.</p>
<p>What got me out of that hole was not prior ostree knowledge, it was looking for a piece of evidence that would settle the question instead of betting on a hunch.</p>
<h2 id="where-this-firmware-comes-from-anyway">Where this firmware comes from anyway</h2>
<p>The Bluetooth blob is not in <code>linux-firmware</code>. MR !946 was closed because the project only accepts blobs submitted by the rights holder — it has to come from MediaTek itself. The Wi-Fi one got in through MR !1055, and that is exactly why the Wi-Fi half of the chip works out of the box and the Bluetooth half does not.</p>
<p>You can extract it from ASUS&rsquo;s Windows driver packages with <a href="https://github.com/jetm/mediatek-mt7927-dkms">extract_firmware.py</a> from the <code>jetm/mediatek-mt7927-dkms</code> project. I extracted it from two different packages, independently, and the files come out <strong>byte for byte identical</strong>:</p>
<table>
	<thead>
			<tr>
					<th>Package</th>
					<th>Container</th>
					<th style="text-align: right">Size</th>
			</tr>
	</thead>
	<tbody>
			<tr>
					<td>Bluetooth V1.1147.0.610</td>
					<td><code>mtkbt_v2.dat</code></td>
					<td style="text-align: right">571349 B</td>
			</tr>
			<tr>
					<td>Wi-Fi V5706054</td>
					<td><code>mtkwlan.dat</code></td>
					<td style="text-align: right">571349 B</td>
			</tr>
	</tbody>
</table>
<pre tabindex="0"><code>sha256  2135f2c4220cfa6e8eb9fdf430517098c13b862a95844bbff0153240a768efa8
path    mediatek/mt7927/BT_RAM_CODE_MT6639_2_1_hdr.bin
built   20260611041233
</code></pre><p>Two observations that save time. The path is <code>mediatek/mt7927/</code>, not <code>mediatek/mt6639/</code> — confirm it directly against the module&rsquo;s strings with <code>modinfo -F firmware btmtk</code> instead of inferring it from the error message, because that convention has already moved between kernel versions. And a hash of <code>669c5c99...</code> at roughly 688 KB circulates out there; it did not reproduce here from either package. Two independent extractions agreeing carry more weight than one loose community figure.</p>
<h2 id="what-not-to-do--read-before-reproducing">What not to do — read before reproducing</h2>
<ul>
<li><strong>Do not run <code>modprobe -r btusb</code>.</strong> The MT6639 firmware hangs during a module reload and the device disappears from <code>lsusb</code> persistently. That is how this whole story started. Reboot instead.</li>
<li><strong>Do not install <code>WIFI_*.bin</code> files into a firmware path.</strong> <code>linux-firmware</code> already ships the correct ones, compressed, in <code>/usr/lib/firmware/mediatek/mt7927/</code>. A loose copy silently shadows the newer blob and breaks Wi-Fi that currently works. Only the Bluetooth file should be installed.</li>
<li><strong>Do not expect a reboot to free the controller.</strong> The standby rail keeps it powered. Only a real power cut clears it.</li>
<li><strong>Do not let the loop run once you see the firmware failure.</strong> Every reset cycle risks wedging the chip again and costs another power cycle. Shut down as soon as <code>Failed to load firmware file (-2)</code> shows up.</li>
</ul>
<h2 id="what-should-change-in-the-kernel">What should change in the kernel</h2>
<p>A firmware file that does not exist will still not exist on the next attempt. Retrying the same request thousands of times cannot possibly succeed — and here it does active harm, because it pushes the controller into a state that survives reboots and needs physical intervention.</p>
<p>A retry limit, or a backoff, or simply not retrying on <code>-2</code> when the previous attempt failed for the same reason, would turn this from a wedged device into a log line saying the firmware is missing. I am taking this to <code>linux-bluetooth</code>.</p>
<h2 id="conclusion">Conclusion</h2>
<p>What I take away from this investigation is not the final command, which fits in two lines. It is two other things.</p>
<p>The first is that absence of evidence is not evidence of absence, and a silent tool lies. An empty <code>grep</code> convinced me for quite a while that there was no log, when in fact there was no permission. It is always worth confirming that the tool is actually handing you data before you interpret its silence.</p>
<p>The second is that the second scare — the apparently empty staged <code>/etc</code> — resolved because I went looking for evidence that would settle the question instead of trusting what I thought ostree did. That <code>fstab</code> in the list was worth more than any certainty of mine about how the system works.</p>
<p>And in the end, the missing firmware really was the problem. It just did not break anything on its own: the driver did, trying with infinite persistence to fix something that could not be fixed that way.</p>
<p>If you have this chip and landed here wondering why your Bluetooth does not work, I hope this article saves you the days it cost me. Any questions or corrections, just reach out.</p>
<p>See you next time!</p>
]]></content:encoded></item></channel></rss>