Replies: 5 comments
|
Update — bisected to two specific core animations (they account for essentially all of the ~145 layouts/s) Follow-up measurement on the same machine (0.1.5-rc.2, Windows, Electron desktop shell), after updating to the newest maid-atelier build (which removed the per-frame JS writes): Baseline with the skin's decorative animations running: main thread 60.4%, Pausing animations one group at a time via
So the two core animations dominate:
Hypothesis: these two effects animate a paint-heavy property (or an SVG / pseudo-element child), so every frame re-rasterizes. Making them compositor-friendly (transform/opacity on a promoted layer, or discrete/steps keyframes) should recover most of those ~50 points without a visible change. Repro on a session with an active turn: document.getAnimations()
.filter(a => /sweep|shimmer/.test(a.animationName || ''))
.forEach(a => a.pause()); // main thread 60% → 2%, layouts 1443 → 10 |
|
Update 2 — the resize stutter is a separate mechanism: ResizeObserver style writes × expensive restyle (The animation issue in Update 1 is idle-time cost. This one is the interactive resize path — and it stutters with or without animations running, matching the user's report.) Measured with a real window resize (Win32
Per-sweep cost breakdown (20 steps):
Who writes styles during resize (captured via a
Single-factor isolation, same sweep:
No single factor fixes it: with the skin disabled the restyle cost drops to 2.4 ms/call, but the call count triples (130 → 311), so the main thread still sits at ~90%. Hypothesis: each ResizeObserver write invalidates style/layout for a subtree, and Blink recomputes far more than the node that changed (this profile has 178 stylesheets / ~8.7k rules). N writes per step × expensive recalc saturates the main thread, so the window edge lags behind the mouse. Batching those writes (one write per frame, or setting CSS variables on a common ancestor / using container queries) or otherwise narrowing the invalidation should recover most of it. Repro: |
|
补一条源码级的定位与修复,供上面那条 7 处,两个成因。 布局路径 —— 5 个运行行扫光,都是绝对定位的 300px 光带动画
绘制路径 —— 2 个状态微光,
修复:分支
实测一:真机 A/B(我自己的 0.1.5-rc.1 Web GUI,把补丁直接打在本机安装的 client bundle 上)。用页面里真实注入的 CSS 选择器拼出 300 个 running 行,探针挂在真实页面上,每档取 5 秒窗口:
把动画暂停在固定进度、读真机 CSS 上 实测二:仓库自己的关卡。 另外把 三点保留:
一个复现细节:lightningcss 会给动画名加 CSS 模块哈希,真机上看到的是 |
|
感谢这么彻底的定位和修复—— 环境:Windows 11 + Electron 桌面壳 + 0.1.5-rc.2 + 第三方皮肤 maid-atelier。 A/B 一(空会话页——页面无工具行,因此没有覆盖你们那 5 处
A/B 二(有工具行的自然会话,0.1.5-rc.2):空闲主线程 47%–57%;把 better-display 的两处 你们说的 5 处扫光我们没能单独测到(空页触发不了),但从你们 300 行探针的 107.6 layouts/s / 1607ms 看,和我们自然会话里剩余的那部分量级一致——这大概是我们这边「用起来卡」的最大一块。 一个来自第三方皮肤侧的旁证:maid-atelier 早期也踩过同类坑(per-frame 写内联样式 + 强制布局,空闲主线程约 51%);作者「削减每帧写入」修完后,我们实测降到 3.3%(1453 layouts/10s → 2)。所以「动画只碰合成层属性、不写样式」这条思路对第三方皮肤同样成立。 等 |
|
谢谢复测,两条数据都很有用——尤其 A/B 二,这是我看到的第一份把这处 1. 2. 空会话也能确定性触发那 5 处扫光(不用等 release)。页面里的样式表已经是真实编译产物,从里面取出选择器、拼出运行行即可: const css = [...document.querySelectorAll('style')].map(s => s.textContent).join('\n')
const rules = [...css.matchAll(/([^{}]+)\{[^{}]*animation:2\.6s ease-out infinite [A-Za-z0-9_-]*row-sweep[^{}]*\}/g)].map(m => m[1].trim())
const mk = s => { const el = document.createElement('div')
for (const c of s.match(/\.[A-Za-z0-9_-]+/g) ?? []) el.classList.add(c.slice(1))
for (const a of s.match(/\[[^\]]+\]/g) ?? []) { const [k, v] = a.slice(1, -1).split('='); el.setAttribute(k, v ?? '') }
return el }
const host = document.createElement('div'); host.id = 'sweep-probe'
host.style.cssText = 'position:fixed;left:0;top:0;width:720px;z-index:99999'
for (let i = 0; i < 60; i++) for (const sel of rules) {
const parts = sel.replace(/:after$/, '').trim().split(/\s+/).map(mk)
if (parts[1]) parts[0].appendChild(parts[1]); host.appendChild(parts[0]) }
document.body.appendChild(host)
// 之后照常用 CDP 的 Performance.getMetrics 取 LayoutCount / TaskDuration;收尾 host.remove()这段会建 300 个运行行(5 个选择器 × 60)。在 0.1.5-rc.1 上是 107.6 layouts/s、5 秒 1607ms 主线程任务;和「移除探针」的基线各测一次,就能把 5 处扫光单独摘出来。一个细节:哈希后的动画名只出现在样式表里, 3. maid-atelier 那条旁证很有价值:1453 → 2 layouts/10s 说明「动画只碰合成层属性、不写样式」对宿主和第三方都成立。等分支进 release 欢迎按老办法复测;先给个预期——微光那两处仍是每秒 60 次 style recalc,修完后不会回到「全暂停」的基线,应落在两者之间。 |
Uh oh!
There was an error while loading. Please reload this page.
环境
0.1.5-rc.2(Windows 桌面 Electron 封装,内核同 Web UI)GpuPreference=2强制独显;glRenderer = ANGLE (NVIDIA …))gpu_compositing/rasterization/webgl均为 enabledforce-prefers-reduced-motion,document.getAnimations()无运行中动画@linxin666/dsh-web-all、@linxin666/dsh-client-ui-skin-center、@anionex/dsh-turn-rewind、@dsh-external/dsh-client-ui-skin-*、dsh-meme、dshmarket等 13 个前端插件)现象
实测数据(CDP 采集,均为无注入的干净测量)
①
Performance.getMetrics,10 秒窗口布局在 10 秒内均匀分布 ≈ 145 次/秒,即几乎每一屏幕帧都在布局。
② trace(
devtools.timeline,6 秒):Layout157 次、UpdateLayoutTree143 次、Paint113 次;无BeginMainFrame/DrawFrame事件。rAF 回调仅 144 次 / 6 秒(24 次/秒)——布局频率(145/s)远高于 rAF 频率(24/s),说明布局不是 rAF 驱动的。③ CPU Profile(8 秒):
(idle)4.5 s、(program)3.5 s(原生渲染工作,非 JS);JS 侧热点toBottom79.8 ms、getBoundingClientRect58.6 ms、closest11.6 ms。④ DOM 变更与强制布局:MutationObserver 采样约 30 次/秒 DOM 变更(多为
data-chat-*、data-dsh-part、hidden属性切换);hookgetBoundingClientRect测到约 59 次/秒。已排除
<style>表 10 秒,随后恢复):主线程 51.2% → 29.0%。皮肤约占一半开销,但禁用后仍有 29% 的持续占用,故皮肤不是唯一原因。猜测与提问
getBoundingClientRect读取 —— 怀疑存在**每帧强制同步布局(layout thrashing)**或 ResizeObserver 反馈循环。是否与长会话消息折叠使用hidden="until-found"/content-visibility有关?(参见 [建议/Web体验] 建议优先保证前端流畅度而非 hidden="until-found":折叠工具与思考时真正卸载 DOM,解决长会话严重卡顿 #6156)相关讨论(先读过,避免重复)
hidden="until-found")backdrop-filter)附:本机第三方插件的额外开销(供参考,非本报告主题,便于复现时对照)
@linxin666/dsh-web-all的皮肤标记引擎applyToTreequerySelectorAll→ 实测 634 次/秒syncHeightgetBoundingClientRectdsh-web-all内置 skin-center + 独立@linxin666/dsh-client-ui-skin-centerAll reactions