<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Zeurd 的博客</title>
        <link>https://www.zeurd.com/</link>
        <description>一个面向 AI / 算法 / 视觉系统 / 工程实践的研究笔记站。</description>
        <lastBuildDate>Thu, 16 Jul 2026 12:49:17 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>zh-CN</language>
        <copyright>All rights reserved 2026, Zeurd</copyright>
        <item>
            <title><![CDATA[科技抱团掩盖了熊市真相]]></title>
            <link>https://www.zeurd.com/article/tech-crowding-hides-the-bear-market</link>
            <guid>https://www.zeurd.com/article/tech-crowding-hides-the-bear-market</guid>
            <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[指数被 AI 与科技权重托住，并不代表市场已经进入牛市。真正值得观察的，是板块轮动、赚钱效应，以及传统消费资产何时重新获得宏观支撑。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39f37eaec0b181fab7edf617cc9a0be8"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><div class="notion-text notion-block-e088f5126caf4270bbc5b95b42ebb8e0">刚刚和朋友聊到老登股，尤其是茅台。</div><div class="notion-text notion-block-092d9cd0dc664bd39d5e150a3b693ed5">说实话，虽然我一直待在科技行业，也主要关注科技股，但我并不看空茅台。因为茅台的核心逻辑，本来就不只是白酒。</div><div class="notion-text notion-block-8902535517e34b14b197e1c18186f96f">它更像是奢侈品和消费品的结合体。茅台不需要和其他白酒比较，就像爱马仕也不会天天和普通箱包讨论性价比。真正支撑它的，是品牌、稀缺性和社会共识。</div><div class="notion-text notion-block-151df2ec11464fe59f72d40123fd056f">而且酒比包还有一个很特殊的地方。</div><div class="notion-text notion-block-ee32e1521e5c4cbb9258b74091c1f5a6">买了一只爱马仕，只要不是狂热收藏，短时间内通常不会再买一只，旧包也不会因为用过一次就消失。但茅台喝完就没了，实物直接归零，只要消费场景还在，需求就可以不断重复。</div><div class="notion-text notion-block-0ab4d823795a436db4002a0d7589a0ce">更现实一点，送爱马仕很容易被认定为明确的利益输送；两箱茅台放在饭局上，最后没喝完顺手塞进后备箱，在很多场景里甚至只会被说成铺张浪费。它既有奢侈品的身份属性，又保留了消费品的模糊边界，这也是茅台长期强势的一部分原因。</div><div class="notion-text notion-block-89a4985d7001476a93295d4344d2eda2">所以我看好茅台的商业逻辑，但这不代表我认为它现在就应该上涨。</div><div class="notion-text notion-block-ef1212b7e57f4c25ad38cd601c754854">我一直有一个比较坚定的判断：现在仍然是熊市。</div><div class="notion-text notion-block-7ca9f76d3a9c405fbeb6f587300e77bb">经济放开以后没有如期回暖，本身就是很明显的信号。只是这一轮 AI 的产业逻辑足够硬，又碰上全球科技共振，资金开始极致抱团。科技股权重又高，少数大票不断上涨，就把指数托住了。</div><div class="notion-text notion-block-312368e25e194ed1b67e293849d4d647">指数没崩，看起来就不像熊市。但把指数拿掉以后，现在的市场状态其实非常熟悉。</div><div class="notion-text notion-block-0e4f43d397d648e7b3f4609c69082665">过去我们怎么理解牛市？</div><div class="notion-text notion-block-08c1d533e6ee45389d363d86f9a7e3b8">牛市不是只有几个方向涨，而是板块不断轮动。最开始涨确定性最高的资产，随后资金向其他行业扩散，最后连最差的股票也能补涨。直到市场找不到新的东西可涨，估值和情绪一起走到尽头，行情才会崩掉。</div><div class="notion-text notion-block-3a41c8646b3b43f197aa9d6b21360627">熊市正好相反。大多数股票持续下跌，资金不愿意承担扩散风险，只能集中到少数仍然有故事、有业绩或者有强烈预期的标的上。所谓熊市多妖股，本质就是缺少普遍机会以后，存量资金在少数方向里极致抱团。</div><div class="notion-text notion-block-32b47117c7fa4b85927f791f586a8cd1">现在除了指数没有明显崩掉，其他表现不就是熊市吗？</div><div class="notion-text notion-block-9728e7657e2f436599028de5e4c82681">大部分行业没有持续赚钱效应，板块轮动很弱，资金高度集中在 AI 和科技，传统消费资产则不断承受经济预期下修。指数上涨更多是权重结构带来的结果，而不是经济和企业盈利普遍改善。</div><div class="notion-text notion-block-4eb5134871ba4f5eb9127f0ff27e80b0">所以茅台真正要重新走强，靠的不会是某一次反弹，也不是资金忽然觉得它便宜了，而是居民收入、消费意愿和商务活动重新繁荣起来。</div><div class="notion-text notion-block-3baf57eea6e549399b7737a053be7c7b">它需要等的是真正的经济繁荣，而不是现在这种由 AI 产业繁荣托起来的指数繁荣。</div><div class="notion-blank notion-block-39f37eaec0b18053bb02c7200fdebd68"> </div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[给 VLM 加一个标记 token，顺便补了下 tokenizer]]></title>
            <link>https://www.zeurd.com/article/vlm-tokenizer-special-token-notes</link>
            <guid>https://www.zeurd.com/article/vlm-tokenizer-special-token-notes</guid>
            <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[因为要给 VLM 新增一个标记 token，重新梳理 BPE、Unigram、词表大小、embedding tying、packing，以及新增 token 在 LoRA 训练里的处理。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39e37eaec0b1810db2dcc049229cc5bb"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><div class="notion-text notion-block-b860f66c29824b2a8cc66047f7f8faa7">最近要给我们自己的 VLM 加一个标记 token。以前一直直接用 Qwen 现成的 tokenizer，从来没认真看过分词表怎么来的。真正要改才发现，新增 token 这件事本身不复杂，后面还连着 embedding、lm_head、loss mask、LoRA 和生成停止条件，漏一个就有可能白训。</div><div class="notion-text notion-block-497e1e66807e4c539a4fe82a210757f2">顺便把 tokenizer 相关的东西重新补了一遍，记在这里，免得以后又从头查。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-923f1019662d4bc5ab13848445c5db47" data-id="923f1019662d4bc5ab13848445c5db47"><span><div id="923f1019662d4bc5ab13848445c5db47" class="notion-header-anchor"></div><a class="notion-hash-link" href="#923f1019662d4bc5ab13848445c5db47" title="BPE 和 Unigram"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">BPE 和 Unigram</span></span></h3><div class="notion-text notion-block-b5a1db7cf2f946899944f5ce2a5eabd1">常见的子词算法主要有 BPE 和 Unigram。</div><div class="notion-text notion-block-895ebfff98e149449bf712ab404a26ab">BPE 从字符或 byte 这类小粒度单位开始，统计训练语料里最常出现的相邻 token 对，每轮合并一对，直到词表达到预设大小。它学到的是一套有顺序的 merge rules，实际分词时按照这套规则执行。</div><div class="notion-text notion-block-1b7bf5e1ef1246b3912745e93138cb64">Unigram 的方向相反。它先从训练语料构造一个较大的候选词表，给每个 token 一个概率，再反复删除对整体似然贡献较低的候选，最后保留到目标词表大小。分词时可以比较同一句话的多种切法，选择概率最高的结果；训练阶段也可以采样不同切法做 subword regularization。</div><div class="notion-text notion-block-1b69e02a564c4018a7282a52a6cc2e54">我最开始看到“从大候选词表开始”，第一反应是它需要事先穷尽所有大词，那使用范围应该很有限。这个理解不对。候选词表通常从训练语料里的高频子串生成，并不需要枚举语言中所有可能的词；常见字符、<code class="notion-inline-code">&lt;unk&gt;</code> 或 byte fallback 还可以负责兜底。Unigram 的麻烦主要在候选集构造、概率估计和反复裁剪，算法与实现都比贪心合并更复杂。</div><div class="notion-text notion-block-bd258bc7779849c99d8b38833ecf7dbc">现在很多 decoder-only LLM 使用 BPE，尤其常见 Byte-level BPE。Unigram 仍然用在不少 SentencePiece、机器翻译和多语言模型里，不能简单概括成已经被 Google 放弃。</div><div class="notion-text notion-block-4694c66ccebc487b8dd971cf82a04b52">这里说的“训练 tokenizer”也是传统算法训练，主要是频率统计、概率估计、动态规划和贪心合并，和训练神经网络没有关系。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-dcf17e3c7faf4ff4b4af5ffe98ecb685" data-id="dcf17e3c7faf4ff4b4af5ffe98ecb685"><span><div id="dcf17e3c7faf4ff4b4af5ffe98ecb685" class="notion-header-anchor"></div><a class="notion-hash-link" href="#dcf17e3c7faf4ff4b4af5ffe98ecb685" title="tiktoken 和 SentencePiece"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">tiktoken 和 SentencePiece</span></span></h3><div class="notion-text notion-block-91d9ed0f9b8449c99f2a752230f999b5">tiktoken 和 SentencePiece 是工具，BPE 和 Unigram 才是算法。</div><div class="notion-text notion-block-da02cf4776884cefa53f43264e7e6f5b">SentencePiece 同时支持 BPE 和 Unigram。它直接处理原始 Unicode 文本，并把空格转成特殊字符 <code class="notion-inline-code">▁</code>，所以经常会看到 <code class="notion-inline-code">▁hello</code> 这种 token。开启 byte fallback 后，词表覆盖不到的字符还可以继续退化成 UTF-8 bytes。</div><div class="notion-text notion-block-964d6c3a3ab34dd1836a399ce2912271">tiktoken 是 OpenAI 的高性能 Byte-level BPE 实现。文本先按 UTF-8 转成 bytes，最底层有 0～255 共 256 个 byte，理论上任意文本都可以表示。训练之后，常见 byte 序列会逐步合并成更长的 token。</div><div class="notion-text notion-block-48c398592ee14ad3aeeb215fb069329e">tiktoken 在执行 BPE 前还会用正则做 pre-tokenization，先把字母、数字、标点和换行等内容拆成片段。后面的 BPE merge 通常不会跨过这些片段边界。</div><div class="notion-text notion-block-b9b39c4e462f4f38ab0ecdec08012a65">所以“常规 BPE 禁止多词组合”这个说法不够准确。限制多半来自 pre-tokenizer，BPE 本身只关心相邻符号能不能合并。SuperBPE 做的事情就是调整这层边界。第一阶段保留常规词边界，先学习 subword；第二阶段放开边界，继续合并高频的多词组合。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-965a99c09d894955a586de1fb854b5a2" data-id="965a99c09d894955a586de1fb854b5a2"><span><div id="965a99c09d894955a586de1fb854b5a2" class="notion-header-anchor"></div><a class="notion-hash-link" href="#965a99c09d894955a586de1fb854b5a2" title="中文是怎么切的"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">中文是怎么切的</span></span></h3><div class="notion-text notion-block-c2246bf2de2e4f1a818ed6d14beb3a8c">英文有天然空格，中文没有。中文 tokenizer 一般会形成几层不同粒度的 token：</div><ul class="notion-list notion-list-disc notion-block-15f2c30c164c4cb5aa04f1f7c0917b00"><li>最底层的 UTF-8 bytes，用来给生僻字和异常字符兜底；</li></ul><ul class="notion-list notion-list-disc notion-block-6528f29488224e06a238ae28f1d3537b"><li>单个汉字，例如“中”“国”“人”“工”“智”“能”；</li></ul><ul class="notion-list notion-list-disc notion-block-d46b3c67ae484784accd4d31d40aabd9"><li>常见词，例如“中国”“人工”“智能”“模型”；</li></ul><ul class="notion-list notion-list-disc notion-block-53e126bd09d745309df7a61a9ccf539c"><li>更长的高频组合，例如“人工智能”“大语言模型”。</li></ul><div class="notion-text notion-block-b8836cbd5ffd48dfbc1efca298b1ffe9">Byte-level BPE 从 bytes 开始。以“中”为例，它的 UTF-8 是 <code class="notion-inline-code">E4 B8 AD</code>，训练时可能先把这三个 byte 合并成“中”，再继续得到“中国”“人工智能”等更长 token。Qwen 这类 tokenizer 走的主要就是这条路线。</div><div class="notion-text notion-block-215cc34d1d424623aee5fc51a190435e">它的好处是编码统一、没有 OOV，也适合中英和代码混合。代价是词表里可能存在落在 Unicode 字符内部的 byte token。完整 token 序列仍然可以无损还原，但单独查看某个 token 时可能出现乱码，词表也没那么容易解释。</div><div class="notion-text notion-block-f1522bfbad054935b12050ffc1645e8c">另一条路线是 Unicode SentencePiece BPE 或 Unigram，再配 byte fallback。常见汉字直接作为基础字符学习，多字组合继续合并，覆盖不到的字符才退回 bytes。这类词表通常更直观，也更容易避免把常见汉字拆在 UTF-8 中间。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-c161eccb53024667abe5c1148d3ec859" data-id="c161eccb53024667abe5c1148d3ec859"><span><div id="c161eccb53024667abe5c1148d3ec859" class="notion-header-anchor"></div><a class="notion-hash-link" href="#c161eccb53024667abe5c1148d3ec859" title="词表大小"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">词表大小</span></span></h3><div class="notion-text notion-block-785a35b386fc4d3cb21dc6d4a3b5bace">词表大小记作 <code class="notion-inline-code">V</code>，模型隐藏维度记作 <code class="notion-inline-code">d</code>，输入 embedding 的参数量就是 <code class="notion-inline-code">V × d</code>。</div><div class="notion-text notion-block-a4e92a2f462a4815828738aa4de93fea">词表变大以后，同一段文本通常可以压缩成更少的 token，序列更短，attention 和 KV Cache 的压力也会下降。与此同时，embedding 和输出层会变大，输出 softmax 更贵，低频 token 分到的训练次数也更少。</div><div class="notion-text notion-block-4445f66f66bd46e58814c1b06bd20d9c">词表较小时，embedding 占用更低，单个 token 的出现次数通常更多，文本却会被拆成更长的序列，训练和推理需要处理更多位置。</div><div class="notion-text notion-block-b4925278f6364614b7a06ad25a70bbdb">所以词表大小同时影响参数量、文本压缩率、多语言覆盖和推理成本。</div><div class="notion-text notion-block-acb3313687c5464ea38a56fb53f403a7">Qwen3.5-2B 的公开配置里，词表大小是 248,320，文本隐藏维度是 2,048。单个 embedding 矩阵有：</div><div class="notion-text notion-block-aa76ce5fcc384188ab1fdf54be428735"><code class="notion-inline-code">248,320 × 2,048 = 508,559,360</code></div><div class="notion-text notion-block-1097a354902f4fe08b5d687d31cfb1ae">也就是约 5.09 亿参数。Qwen3.5-2B 配置了 embedding tying，输入 embedding 和 lm_head 共用这一个矩阵。没有 tying 的话，输出层还要再放一份同样大小的权重。tying 省的是参数量和模型存储，<code class="notion-inline-code">V × d</code> 的输出投影计算仍然要做。</div><div class="notion-text notion-block-a1e47aefeed4411582b31144ad2a3806">Token fertility 通常表示一段文本平均每个词会产生多少 token。英文可以按 word 统计，中文的“词”边界不统一，也经常按字符或固定语料的 token 数直接比较。fertility 越高，同样的内容需要的 token 越多，能放进上下文的有效内容越少，训练和推理成本也会跟着增加。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-48a7fb2dfe8440e2bafa9962d3d866d1" data-id="48a7fb2dfe8440e2bafa9962d3d866d1"><span><div id="48a7fb2dfe8440e2bafa9962d3d866d1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#48a7fb2dfe8440e2bafa9962d3d866d1" title="新增一个标记 token"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">新增一个标记 token</span></span></h3><div class="notion-text notion-block-81727c3b3ca84063b75addd36564cc36">先要区分它只是一个希望保持完整的普通 token，还是模型协议里的控制 token。</div><div class="notion-text notion-block-ceca95ad88cb482590b12115d96e2bfd">普通领域词只需要保证不被拆开，可以用 <code class="notion-inline-code">add_tokens</code>。像角色边界、模态占位符、输出标签这类控制符，再放进 <code class="notion-inline-code">additional_special_tokens</code>。special token 会影响 tokenizer 的特殊 token 处理逻辑，不需要把所有新增词都设成 special。</div><div class="notion-text notion-block-1017c0e890544c6397d2903e24b10264">用 Transformers 增加一个控制 token，大致是这样：</div><div class="notion-text notion-block-072415bdfe6e473bb36a75561277c2c4"><code class="notion-inline-code">resize_token_embeddings</code> 会扩展输入 embedding。模型类实现了 <code class="notion-inline-code">tie_weights()</code> 时，Transformers 也会重新处理 tied weight，一般不用自己手动绑定 token ID、embedding 和 lm_head。</div><div class="notion-text notion-block-ea64cc9887904be5a66a1902d3b08618">后面还有几件事要检查：</div><ol start="1" class="notion-list notion-list-numbered notion-block-24f3b90c0cf04aaaa83e1a8fd765ac80" style="list-style-type:decimal"><li>新 token 是正常输入时，attention mask 设为 1。只有 padding 位置才设为 0。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-c37450a1f0244963b1438ae74028758e" style="list-style-type:decimal"><li>Loss mask 取决于模型要不要学会生成它。只作为输入提示或 padding 时可以设为 <code class="notion-inline-code">-100</code>；它属于输出协议时要保留 loss。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-5004b8d2750643a4ae84abbf6205a298" style="list-style-type:decimal"><li>只有确实表示生成结束的 token 才加入 <code class="notion-inline-code">eos_token_id</code> 或 stopping criteria。普通分隔符不能一股脑设成 stop token。</li></ol><ol start="4" class="notion-list notion-list-numbered notion-block-63ccfa1345744f7e84bf9b1e98668997" style="list-style-type:decimal"><li>Chat template 只在对话协议跟着变化时修改。</li></ol><ol start="5" class="notion-list notion-list-numbered notion-block-75581e1cb49f43f6a747da2ae8de6a40" style="list-style-type:decimal"><li>tokenizer 和模型要一起保存。只保存 adapter，漏掉改过的 tokenizer，部署时 token ID 很容易对不上。</li></ol><ol start="6" class="notion-list notion-list-numbered notion-block-d6fb1c5ea28246128ec7b760e7e6c816" style="list-style-type:decimal"><li>训练数据里必须真的出现这个 token。扩一行 embedding 不会自动赋予它语义。</li></ol><div class="notion-text notion-block-b8a1419ae24a435199e55c3f00dd52ce">新增 token 也不一定要全参微调。普通 LoRA 默认冻结 embedding 和 lm_head，新 token 的向量确实学不到。PEFT 现在可以用 <code class="notion-inline-code">trainable_token_indices</code> 只训练新增 token 对应的 embedding 行；也可以把 embedding 或 lm_head 放进 <code class="notion-inline-code">modules_to_save</code>，代价是对应模块会整体保存和训练。</div><div class="notion-text notion-block-5a8bfb11999f4c7981ddddd1d91c6e77">这里还要看模型有没有 embedding tying。tied 模型的输入和输出共用权重；untied 模型里，输入 embedding 和 lm_head 是两份参数，只训练输入侧还不足以让模型学会输出新 token。加完 LoRA 后最好直接检查新增行有没有梯度，保存、加载一次 adapter，再确认 token ID 和权重都还在。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-06486bc8da9e46a0ad53f2518fa3286a" data-id="06486bc8da9e46a0ad53f2518fa3286a"><span><div id="06486bc8da9e46a0ad53f2518fa3286a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#06486bc8da9e46a0ad53f2518fa3286a" title="Embedding tying"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Embedding tying</span></span></h3><div class="notion-text notion-block-a88b7a6df89641c589d40ad142a52a74">Embedding tying 就是让同一个 token 的输入表示和输出分类器共用一组向量。它能少掉一个 <code class="notion-inline-code">V × d</code> 矩阵，也给输入空间和输出空间加了一层约束。</div><div class="notion-text notion-block-a4bfa2afa3504f16940fe8797940dc1d">不做 tying 也有合理性。输入表示和输出分类毕竟是两个不同任务，拆开以后两套权重可以各自学习。不同模型两种方案都有，不能只看有没有 tying 判断结构好坏。</div><div class="notion-text notion-block-e2042612195342e5b63c21fc432bd4d3">拿 CV 类比，embedding 有点像网络入口，lm_head 类似最后的分类层。这个类比只能帮助理解它们所在的位置。Embedding 实际是按 token ID 查表，第一层卷积还包含局部空间计算，两者的运算并不一样。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-1639b1cb300843b899878d8abf826bb7" data-id="1639b1cb300843b899878d8abf826bb7"><span><div id="1639b1cb300843b899878d8abf826bb7" class="notion-header-anchor"></div><a class="notion-hash-link" href="#1639b1cb300843b899878d8abf826bb7" title="分词边界会不会影响模型"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">分词边界会不会影响模型</span></span></h3><div class="notion-text notion-block-10b7ef804a5644f294722a6448e5e3a9">我之前还有个疑惑。同一句“今天天气真好”，切成“今天 / 天气 / 真好”，或者“今 / 天天 / 气 / 真好”，模型看到的序列完全不同，结果当然也可能受影响。</div><div class="notion-text notion-block-9d95191d4930474db3c8f0de8927e0d7">固定 tokenizer 对同一段完整文本的切分通常是确定的。训练时“今天天气真好”怎么切，推理时遇到完全相同的字符串仍然会怎么切。换成“今天的天气真好”以后，局部 token 组合可能变化，模型需要靠上下文表示去适应。</div><div class="notion-text notion-block-ba94839362ab45c2903bf6a870fcae34">token 本来也不要求和人理解的词义一一对应。“今天”可以是一个 token，也可以拆成两个；后面的 Transformer 仍然会通过上下文把它们组合起来。切分质量会影响序列长度、低频片段的训练密度和模型对拼写变化的鲁棒性，影响大小要看语料与任务，不能只从某个词切得顺不顺判断。</div><div class="notion-text notion-block-bc6d49addf2a4e1aa9d0a1409136b8ea">Unigram sampling 和 BPE-dropout 还会在训练阶段主动给同一句话采样不同切法，提高模型对分词变化的适应能力。常规 LLM 在推理时一般仍然使用固定切分。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-862833be9b104632a46124b65622e709" data-id="862833be9b104632a46124b65622e709"><span><div id="862833be9b104632a46124b65622e709" class="notion-header-anchor"></div><a class="notion-hash-link" href="#862833be9b104632a46124b65622e709" title="Packing"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">Packing</span></span></h3><div class="notion-text notion-block-5da227e06eef4f27ace699ebd1d19363">训练样本长度不一样，全部 pad 到同一长度会浪费很多计算。Packing 会把多个短样本塞进一个固定长度序列，尽量减少 padding。</div><div class="notion-text notion-block-18c9a0a53288457fbba3fda508d8706e">实现时不能只做简单拼接。样本之间至少要放 EOS 或其他边界标记，还要处理好 labels、position IDs 和 attention 边界。SFT 里通常只给 assistant 输出计算 loss；是否允许不同样本互相 attention，也要和训练框架的实现对齐。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-2f3030a0e681451b934659d7ebabf2f8" data-id="2f3030a0e681451b934659d7ebabf2f8"><span><div id="2f3030a0e681451b934659d7ebabf2f8" class="notion-header-anchor"></div><a class="notion-hash-link" href="#2f3030a0e681451b934659d7ebabf2f8" title="相关资料"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">相关资料</span></span></h3><ul class="notion-list notion-list-disc notion-block-3c15514e73e4477da06d61711a6f0fd4"><li><a class="notion-link" href="https://github.com/google/sentencepiece" target="_blank" rel="noopener noreferrer">SentencePiece</a></li></ul><ul class="notion-list notion-list-disc notion-block-7d641a84faf7476da1414cb550de1308"><li><a class="notion-link" href="https://github.com/openai/tiktoken" target="_blank" rel="noopener noreferrer">tiktoken</a></li></ul><ul class="notion-list notion-list-disc notion-block-c80fe64cd3a842cd91606358010edf83"><li><a class="notion-link" href="https://arxiv.org/abs/2503.13423" target="_blank" rel="noopener noreferrer">SuperBPE: Space Travel for Language Models</a></li></ul><ul class="notion-list notion-list-disc notion-block-03f47baa565d43839103d2d7a9d77b0c"><li><a class="notion-link" href="https://huggingface.co/Qwen/Qwen3.5-2B/blob/main/config.json" target="_blank" rel="noopener noreferrer">Qwen3.5-2B config</a></li></ul><ul class="notion-list notion-list-disc notion-block-f6fbbadd35094633a4f8dfdf2ef3962d"><li><a class="notion-link" href="https://huggingface.co/docs/transformers/main_classes/model" target="_blank" rel="noopener noreferrer">Transformers: resize_token_embeddings</a></li></ul><ul class="notion-list notion-list-disc notion-block-8db36fe9bb094f0a8a2841f9a306fba8"><li><a class="notion-link" href="https://huggingface.co/docs/peft/package_reference/lora" target="_blank" rel="noopener noreferrer">PEFT LoRA: trainable_token_indices</a></li></ul></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[光流估计论文阅读小结]]></title>
            <link>https://www.zeurd.com/article/zhihu-1905637782333928584</link>
            <guid>https://www.zeurd.com/article/zhihu-1905637782333928584</guid>
            <pubDate>Tue, 13 May 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[从 FlowNet 到 IRR-PWC，记录六篇光流论文的核心结构和关键消融，补充训练数据、参数量与 benchmark 对比。]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b1811fbd72c0062f060cce"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181b8b16ce4c137956665"/><div class="notion-text notion-block-39737eaec0b181eb84c2f4cf38a71551">最近在研究光流估计相关的论文工作，在此记录一下光流估计的发展历程。从最简单的 CNN 换个任务，到后续这么多年的发展，希望能梳理出一个脉络。也希望相关方向的大佬们不吝赐教，多多交流。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-31b632fd03694768bfde2b376ff4c22a" data-id="31b632fd03694768bfde2b376ff4c22a"><span><div id="31b632fd03694768bfde2b376ff4c22a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#31b632fd03694768bfde2b376ff4c22a" title="传说之始"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">传说之始</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181158894fa79ca689a8b" data-id="39737eaec0b181158894fa79ca689a8b"><span><div id="39737eaec0b181158894fa79ca689a8b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181158894fa79ca689a8b" title="FlowNet: Learning Optical Flow with Convolutional Networks【ICCV 2015】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default"><a class="notion-link" href="https://arxiv.org/abs/1504.06852" target="_blank" rel="noopener noreferrer">FlowNet: Learning Optical Flow with Convolutional Networks</a></span><span class="notion-default">【ICCV 2015】</span></span></span></h4><div class="notion-text notion-block-39737eaec0b1814593b0f4cf92749a70">传说，在最早的时候，CNN 能做的任务中并不包含光流估计。直到 FlowNet 给出第一套端到端监督 CNN，为 CNN 的传说又画上了浓墨重彩的一笔。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b181afa9ddc1342cf56bc1"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic2.zhimg.com/v2-c7d729d454abe57e093be0bea75f76bb_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-81af-a9dd-c1342cf56bc1" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181d3acbdd19070bae67d">作者做了两套结构。FlowNetS 直接把两帧图像 cat 到一起，FlowNetC 用双分支提取特征，再通过 correlation layer 做匹配。这个 correlation 没有可学习参数，本质上就是两个位置特征向量的点积。decoder 会把上采样特征、encoder 同尺度特征和上一层 flow 拼起来，最终学习到四分之一分辨率光流，再插值回原图。</div><div class="notion-text notion-block-1ad31a259151474dbf996071c69d5819">训练数据 FlyingChairs 一共有 22872 对图像，其中 22232 对训练，640 对测试。它用 Flickr 背景和椅子图片合成二维运动，画面很假，却比直接拿小规模 Sintel 训练有效。论文里，先用 Chairs 预训练再到 Sintel fine-tune，大约能少 1 pixel EPE；拿掉数据增强，Sintel 上还会再差约 2 pixel。</div><div class="notion-text notion-block-23a59519c8d8411b922071599af2b2a4">FlowNetC 也没有稳稳压过 FlowNetS。FlyingChairs test 上两者是 2.19 和 2.71，到了 Sintel Final test，FlowNetC 是 8.81，FlowNetS 则是 8.43。correlation 在训练域和干净画面上有帮助，碰到 motion blur、fog 或更复杂的运动就不一定了。</div><div class="notion-text notion-block-1a870ed95b17498890858d8f996ef1bd">GTX Titan 上，FlowNetS 和 FlowNetC 分别耗时 80 ms、150 ms。可选的 variational refinement 会把时间增加到 1 秒左右，在 Chairs 上还会让误差变大。FlowNet 证明了端到端 CNN 能做光流，精度仍低于当时最好的传统方法。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181eaa923ca036e8a129e" data-id="39737eaec0b181eaa923ca036e8a129e"><span><div id="39737eaec0b181eaa923ca036e8a129e" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181eaa923ca036e8a129e" title="初露峥嵘"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">初露峥嵘</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b18108ae37f317e7d93351" data-id="39737eaec0b18108ae37f317e7d93351"><span><div id="39737eaec0b18108ae37f317e7d93351" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18108ae37f317e7d93351" title="FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks【CVPR 2017】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><a class="notion-link" href="https://arxiv.org/abs/1612.01925" target="_blank" rel="noopener noreferrer">FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks</a>【CVPR 2017】</span></span></h4><div class="notion-text notion-block-789d2e81f03e4ed18a8c2196a1c6d778">FlowNet 出来之后，问题很快就暴露了。作为开山作，性能却还不如传统方法，小位移和真实场景尤其拉胯。这怎么行，作者赶紧补了一篇 2.0。</div><div class="notion-text notion-block-39737eaec0b181bfaa11db4eb9638522">FlowNet 2.0 主要改了三处。</div><ol start="1" class="notion-list notion-list-numbered notion-block-39737eaec0b181b8b2d1e81b2af0f63a" style="list-style-type:decimal"><li>调整训练数据和训练顺序。</li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39737eaec0b181adbb81f7c4a6bc0fa9" style="list-style-type:decimal"><li>堆叠多个 FlowNet，继续加算力。</li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-490a2dcaf2204f55816ce8036625f42f" style="list-style-type:decimal"><li>增加小位移分支。</li></ol><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181309030d1bb51c7157f" data-id="39737eaec0b181309030d1bb51c7157f"><span><div id="39737eaec0b181309030d1bb51c7157f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181309030d1bb51c7157f" title="1. 改进数据"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">1. 改进数据</span></span></h4><div class="notion-text notion-block-5687b3b0500c4a25a3cf366687816679">作者先在 FlyingChairs 上训练，再用更复杂的 FlyingThings3D 微调。原论文 Table 1 里，FlowNetS 只用 Things3D 训练后在 Sintel Clean train 上是 4.50，两类数据五五混合是 4.10，先 Chairs 再 Things3D 能到 3.79。论文把它叫作 curriculum learning，简单二维运动确实适合拿来做第一阶段训练。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181719cebdb1c62f3ce51" data-id="39737eaec0b181719cebdb1c62f3ce51"><span><div id="39737eaec0b181719cebdb1c62f3ce51" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181719cebdb1c62f3ce51" title="2. 堆叠网络"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">2. 堆叠网络</span></span></h4><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-705ded7495ad471aaa20ae9cb8549261"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic3.zhimg.com/v2-ac5bd77bf698cf0d940c3f52ed15f9cb_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=705ded74-95ad-471a-aa20-ae9cb8549261" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181118247d69f073d7238">第一个网络先预测 flow，用它把第二帧 warp 到第一帧。后面的网络继续接收两张原图、上一版 flow、warped image 和 brightness error，逐级修正结果。</div><div class="notion-text notion-block-39737eaec0b181c5a43be281eb127dc4">Table 2 把 warp 的作用测得很清楚。单个 Net1 在 Chairs test 和 Sintel Clean train 上分别是 3.01、3.79。接第二个网络却不做 warp，Chairs 降到 2.60，Sintel 反而恶化到 4.29。冻结 Net1，再训练带 warp 的 Net2，两个数字变成 1.94、2.93。网络堆得更深也开始收益递减。C、CS、CSS 在 Sintel Clean train 上是 3.04、2.20、2.10，GTX 1080 上耗时则从 33 ms 增加到 51 ms、69 ms。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b1819ab0ecf85bc3449cff" data-id="39737eaec0b1819ab0ecf85bc3449cff"><span><div id="39737eaec0b1819ab0ecf85bc3449cff" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1819ab0ecf85bc3449cff" title="3. 针对小运动"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">3. 针对小运动</span></span></h4><div class="notion-text notion-block-3e66cb75b65e4b239603b3be17d25535">作者又做了 ChairsSDHom，里面大多是小位移和大块平滑区域，再单独训练 FlowNet2-SD。这个分支去掉开头的 stride 2，把 7×7、5×5 大卷积换成多层 3×3，尽量保住小运动。最终模型把大位移分支、小位移分支一起送进 fusion network。</div><div class="notion-text notion-block-39737eaec0b1816eab6ffd5f6b832d63">小运动训练有明确取舍。Middlebury train AEE 从 0.44 降到 0.38，KITTI 2012 train 却从 3.55 变成 4.05。最终 FlowNet2 在 Sintel test 上是 Clean 3.96、Final 6.02，GTX 1080 上耗时 123 ms。误差大幅降下来了，模型也膨胀到了 162M 参数。</div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-49960c9349be44978c505c69b18c6157" data-id="49960c9349be44978c505c69b18c6157"><span><div id="49960c9349be44978c505c69b18c6157" class="notion-header-anchor"></div><a class="notion-hash-link" href="#49960c9349be44978c505c69b18c6157" title="魔童降世"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title">魔童降世</span></span></h3><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-3e69543355fa41708f1edbaec4ee5979" data-id="3e69543355fa41708f1edbaec4ee5979"><span><div id="3e69543355fa41708f1edbaec4ee5979" class="notion-header-anchor"></div><a class="notion-hash-link" href="#3e69543355fa41708f1edbaec4ee5979" title="Optical Flow Estimation Using a Spatial Pyramid Network【CVPR 2017】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><a class="notion-link" href="https://openaccess.thecvf.com/content_cvpr_2017/html/Ranjan_Optical_Flow_Estimation_CVPR_2017_paper.html" target="_blank" rel="noopener noreferrer">Optical Flow Estimation Using a Spatial Pyramid Network</a>【CVPR 2017】</span></span></h4><div class="notion-text notion-block-f4df4fd9465547e0b41ffe863cf0052b">在 FlowNet 茁壮成长的时候，一个小小的灵珠炸弹也在光流估计领域掀起了一点浪花。SPyNet 没有继续堆 FlowNet，走的是传统方法和 CNN 结合的路线。传统光流本来就在做 coarse-to-fine 空间金字塔，放到 deep learning 里同样可以用。</div><div class="notion-text notion-block-39737eaec0b181149d77ce37677ff6ba">顺带说一句，我之前一直以为金字塔是目标检测那边搞出来的东西，这次看光流才知道，传统算法几十年前就在用了。</div><div class="notion-text notion-block-39737eaec0b1817a9a6adb0253d64cc9">它先建立图像金字塔，从最粗的尺度开始估计 flow。到下一层以后，把上一层的 flow 上采样，用它 warp 第二帧，再让当前层的小网络预测 residual flow。大位移在低分辨率上会变小，每层网络只处理剩下的一点残差。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-deb2ac28ec474b2cbbb88ce2d4be35fd"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic4.zhimg.com/v2-502d2cb6ddf177d6e976fb038095cc13_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=deb2ac28-ec47-4b2c-bbb8-8ce2d4be35fd" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b1818ea1bdf04625c1375a">SPyNet 训练了 5 个尺度网络，每个都是独立的小 CNN，包含 5 个 7×7 卷积，各层不共享参数。整套模型只有 120 万参数、9.7 MB，比 FlowNet 小约 96%。论文表里给出的耗时是 69 ms，FlowNetS、FlowNetC 的 80 ms、150 ms 则引用自原论文。几组时间来自不同实验环境，只能看大致量级。</div><div class="notion-text notion-block-950e035c8bf74cc48a18657ba6c14f6e">它也没有在所有 benchmark 上超过 FlowNet。Sintel Clean、Final test 分别是 6.69、8.43；针对 Final fine-tune 后是 8.36，仍高于 FlowNetS+ft 的 7.76。coarse-to-fine 对快速小物体尤其吃亏。Sintel Final 中位移超过 40 pixel 的区域，SPyNet+ft 的 EPE 是 49.71，FlowNetS 和 FlowNetC 分别是 43.24、40.78。小物体在粗尺度消失以后，后面的 residual 网络很难补回来。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-54b2fb02b9174e6f874c634a79dbb40f" data-id="54b2fb02b9174e6f874c634a79dbb40f"><span><div id="54b2fb02b9174e6f874c634a79dbb40f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#54b2fb02b9174e6f874c634a79dbb40f" title="PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume【CVPR 2018】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><a class="notion-link" href="https://openaccess.thecvf.com/content_cvpr_2018/html/Sun_PWC-Net_CNNs_for_CVPR_2018_paper.html" target="_blank" rel="noopener noreferrer">PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume</a>【CVPR 2018】</span></span></h4><div class="notion-text notion-block-aa004194303142b0b0a3fe348f4fda84">SPyNet 打开了传统方法结合 CNN 的大门，一年后 PWC-Net 横空出世。它在空间金字塔上加入可学习特征、feature warp 和局部 cost volume。</div><div class="notion-text notion-block-104f5c483ea24221bc7a711af4b5031c">两帧图像先各自提取特征金字塔。每一层用上一层的 flow 把第二帧特征 warp 到第一帧坐标系，再和第一帧特征计算局部 cost volume，卷积网络据此预测当前层光流。最后还有一个 dilated context network。输出仍是四分之一分辨率，再插值回原图。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-8e140931400f4f939cc10126a6068edb"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pica.zhimg.com/v2-0e74f8a381bade1e4f3bd1e31007ffc7_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=8e140931-400f-4f93-9cc1-0126a6068edb" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-4da0cadc811641e28bae348633f04703">原论文 Table 5 把几个组件逐个拆开。只用 FlyingChairs 训练时，完整模型在 Sintel Clean、Final 上是 3.33、4.59，去掉 warp 后变成 3.79、5.30。context network 也有稳定收益，拿掉以后 Final 从 4.59 变成 4.74。cost volume 的搜索半径倒不用开得很大。每层最大位移设成 2、4、6 时，Final 分别是 4.50、4.59、4.60，金字塔已经把有效搜索范围放大了。</div><div class="notion-text notion-block-eb94069d18e043fb901bd63df3bbc2bc">同一台 Pascal Titan X 上，FlowNet2 有 162.49M 参数，forward 84.80 ms；PWC-Net 是 8.75M 参数和 28.56 ms。PWC-Net-ft 在 Sintel test 上是 Clean 3.86、Final 5.13，KITTI 2015 test 的 Fl-all 是 9.60%。这些是各自 fine-tune 后的模型，不能当成一个模型同时拿到的结果。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-13a3449d23664f91a0b6cece25d42d4b" data-id="13a3449d23664f91a0b6cece25d42d4b"><span><div id="13a3449d23664f91a0b6cece25d42d4b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#13a3449d23664f91a0b6cece25d42d4b" title="LiteFlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation【CVPR 2018】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><a class="notion-link" href="https://openaccess.thecvf.com/content_cvpr_2018/html/Hui_LiteFlowNet_A_Lightweight_CVPR_2018_paper.html" target="_blank" rel="noopener noreferrer">LiteFlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation</a>【CVPR 2018】</span></span></h4><div class="notion-text notion-block-cfc0b51af67040be9460494af05f6103">如果说 FlowNet 2.0 是对 FlowNet 的力大砖飞，那 LiteFlowNet 就是在逐个处理 FlowNet 2.0 里留下的问题。它在每个金字塔层先做 descriptor matching，得到 pixel-level flow，再做 sub-pixel refinement。后面还有一个 Flow Regularization 层，用 feature-driven local convolution 约束局部流场。这里使用 feature warp，FlowNet 2.0 使用 image warp。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-477fa5116f0448fea7fa998a8d99c2c2"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pica.zhimg.com/v2-a00b2e52324d9592a781a9ec9cf03c32_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=477fa511-6f04-48fe-a7fa-998a8d99c2c2" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-83c936f3e20c4f8d8720c3685ee133bc">Table 4 的消融很直接。这些 variant 都只在 FlyingChairs 上训练，再放到 Sintel 和 KITTI 的训练集看迁移。以 Sintel Final 为例，只有 matching 时是 5.69，加 feature warp 后降到 4.81，再加 sub-pixel refinement 是 4.45，最后加 regularization 到 4.17。KITTI 2015 也从 18.24 依次降到 14.52、12.32、11.58，feature warp 带来的那一步最大。</div><div class="notion-text notion-block-d048fd33991a4975b3d33d7a3e2947db">LiteFlowNet 的 “30 倍” 指参数量。相同 GTX 1080 上，它有 5.37M 参数、耗时 90.25 ms；FlowNet2 是 162.49M 和 122.39 ms，速度只快 1.36 倍。论文没有报告 FLOPs 或 MACs。精度也要分数据看。LiteFlowNet-ft 在 Sintel Clean、Final test 上是 4.86、6.09，低于 FlowNet2-ft-Sintel；KITTI 2015 test 的 Fl-all 是 10.24%，又优于 FlowNet2-ft-KITTI 的 11.48%。用 on par 概括更合适。</div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-611db45ef1154d40815b28b0c834bf54" data-id="611db45ef1154d40815b28b0c834bf54"><span><div id="611db45ef1154d40815b28b0c834bf54" class="notion-header-anchor"></div><a class="notion-hash-link" href="#611db45ef1154d40815b28b0c834bf54" title="IRR-PWC: Iterative Residual Refinement for Joint Optical Flow and Occlusion Estimation【CVPR 2019】"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><a class="notion-link" href="https://openaccess.thecvf.com/content_CVPR_2019/html/Hur_Iterative_Residual_Refinement_for_Joint_Optical_Flow_and_Occlusion_Estimation_CVPR_2019_paper.html" target="_blank" rel="noopener noreferrer">IRR-PWC: Iterative Residual Refinement for Joint Optical Flow and Occlusion Estimation</a>【CVPR 2019】</span></span></h4><div class="notion-text notion-block-c6a60246493245a1835a552c6867b278">PWC-Net 之后，IRR-PWC 又把传统优化里的迭代思想搬了进来，同时加入遮挡估计。</div><div class="notion-text notion-block-b4a4ec74cd7e453892bcb01e35d17c40">PWC-Net 原来每个金字塔层都有独立 decoder，IRR-PWC 改成不同层共享同一个 decoder，反复预测 residual 再更新 flow。多跑几轮不会继续增加 decoder 参数。</div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-f7e0f9a4951342f1bc0736b1dc1cf301"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic1.zhimg.com/v2-f721eeec9b08b19dcff75c3d06dbd9be_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=f7e0f9a4-9513-42f1-bc07-36b1dc1cf301" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-0f2f9fd6eb584e0ea52ecb3a15fb1eed">论文先在相同训练设置下比较共享迭代和直接堆网络。跑到第 5 轮时，共享 IRR 在 Sintel Clean train 上是 3.302，独立堆叠 FlowNetS 是 3.517；IRR 从第 3 轮的 3.325 往后已经基本不再提升。PWC-Net 上只加 IRR，Sintel Clean、Final 从 3.13、4.41 降到 2.79、4.10，参数量减少 61.2%。完整模型加入双向估计、遮挡和 refinement 后是 2.34、3.95，参数量仍比 baseline 少 26.4%。</div><div class="notion-text notion-block-b737b8b84e5449ed99c3ea3e972d439a">遮挡分支同时预测前向和反向 flow，再用独立 decoder 预测 occlusion mask。全分辨率 occlusion upsampling 会把 Sintel Final 的 F1 从 0.602 提高到 0.624，flow EPE 反而略有上升。这组实验至少说明上采样模块主要在修 mask，没有顺手把光流也一起抬高。</div><div class="notion-text notion-block-c9a3c1a07bab43558f21bff409b1c890">最终 IRR-PWC 在 Sintel test 上是 Clean 3.84、Final 4.58，参数量 6.36M；相同 fine-tune 设置下的 PWC-Net-ft-final 是 4.39、5.04 和 8.75M。KITTI 2015 test 的 Fl-all 是 7.65%。论文没有报告 runtime、FPS 或 FLOPs，因此只能确认参数量下降，推理时间还没法从论文数字里判断。</div><div class="notion-text notion-block-71af53fdb1f74ec5b128cc081c0d64a6">后面有机会的话，再把 RAFT、GMFlow 这些工作接着补上。</div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[优化][ECCV 2022]Bootstrap Generalization Ability from Loss  Landscape Perspective]]></title>
            <link>https://www.zeurd.com/article/zhihu-719266623</link>
            <guid>https://www.zeurd.com/article/zhihu-719266623</guid>
            <pubDate>Tue, 10 Sep 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[起因是在上班刷知乎的时候看见了 @虚无 大佬的关于一个优化器损失调整方面的回答，感觉还挺有意思的，准备记录一下，下午工作的时候尝试一下。论文名称： Bootstrap Generalization Ability from Loss Landscape Perspective故事会环节在做深度学习的时候，…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b1811db0abdc83fb5ca176"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181e6b3b2f1b88230032a"/><div class="notion-text notion-block-39737eaec0b18152b1b3c3ee2ec5603f"><span class="notion-default">起因是在上班刷知乎的时候看见了 </span><span class="notion-default"><a class="notion-link" href="https://www.zhihu.com/people/d2240cf1c3ac1bd2ada3d2baf0d9efdd" target="_blank" rel="noopener noreferrer">@虚无</a></span><span class="notion-default"> 大佬的关于一个优化器损失调整方面的回答，感觉还挺有意思的，准备记录一下，下午工作的时候尝试一下。</span></div><div class="notion-text notion-block-39737eaec0b18189aacecc31f33783c3"><span class="notion-default">论文名称：</span><span class="notion-default"><a class="notion-link" href="https://arxiv.org/pdf/2209.08473" target="_blank" rel="noopener noreferrer">Bootstrap Generalization Ability from Loss  Landscape Perspective</a></span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18158a72bf7782307c40c" data-id="39737eaec0b18158a72bf7782307c40c"><span><div id="39737eaec0b18158a72bf7782307c40c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18158a72bf7782307c40c" title="故事会环节"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">故事会环节</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181f488b3e4d230138e3b"><span class="notion-default">在做深度学习的时候，会发现一个事儿，就是如果训练数据和测试数据是一致的，那么效果一般都不错，但是如果不一致的时候（这个确实，在实际工业应用中，很少能有一致的情况发生，谁知道用户是用来测啥的）性能就会掉很多，这种情况就被叫做域偏移。常见的解决办法就是做域适应（DA），域泛化（DG）。作者表示DA需要目标域的访问权限，但是DG不用，所以DG被关注的很多，然后就开始说DG。（TIPS：我自己做域相关的比较少，但是就我自己接触的时候，我印象里是DA也很多，特别是UDA，在医学上好像被用的很广泛）DG方案分为3类，数据增强，域不变表示学习，学习策略优化。（这么看确实DG比较多，毕竟基本已经是深度学习默认使用的方案了）</span></div><div class="notion-text notion-block-39737eaec0b181568b01cbb24666a1d1"><span class="notion-default">作者在通过观察后发现，如果训练集和测试集的每个域之间没有什么差异，那么域不变学习不但没用，而且会变差，数据增强和学习策略优化效果很好，然后说这些都可以从landscape的角度解释，帮助模型收敛到平坦最优，取得更好的泛化性。</span></div><div class="notion-text notion-block-39737eaec0b18101ba05d075f6b5346d"><span class="notion-default">然后一堆方法表明平坦的最小值会让泛化性更好。作者于是提出了一系列基于损失的学习策略，来让模型收敛到平坦值。其中包括了：（i）提出一种蒸馏微调范式（ii）用大学习率，然后用一种新的学习率调度器，ALRS来帮助收敛。（iii）把模型加宽得到改进模型（这里其实我有一些，小疑问的，我其实很少看这种方法很多的论文，因为很清楚的知道，一般来说，有好几个改进的原因，都是怕审稿人觉得方法太简单，不够novelty，所以往往会有一两个凑数的方案，不知道本文有没有这种内容）</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181b0b679d01664d2a5e0" data-id="39737eaec0b181b0b679d01664d2a5e0"><span><div id="39737eaec0b181b0b679d01664d2a5e0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181b0b679d01664d2a5e0" title="创新点"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">创新点</span></span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-39737eaec0b1811491a0c148c6438b13" style="list-style-type:decimal"><li><span class="notion-default">论文从loss landscape的角度解释了DG方法work的原因</span></li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39737eaec0b181998982f33ee27c801b" style="list-style-type:decimal"><li><span class="notion-default">提出了一系列方法</span></li></ol><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181aca45ccef14d7bb1bb" data-id="39737eaec0b181aca45ccef14d7bb1bb"><span><div id="39737eaec0b181aca45ccef14d7bb1bb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181aca45ccef14d7bb1bb" title="Method"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Method</span></span></span></h3><div class="notion-text notion-block-39737eaec0b1810f8094d956fe2b2a6f"><span class="notion-default">本文所谓的蒸馏微调范式，总共4步：1-正常训练一个分类模型。2-使用自蒸馏蒸一下。3-把分辨率调大再训一个分类模型。4-再自蒸馏蒸一下。每次训练都是从头开始训，学习率重置.(在22年的时候不清楚，现在的话，学习率慢慢方法其实是一个通用trick，关注挑战赛的应该能发现，十几个工作有大半都会用这个方式，同时经常用的还有例如，多teacher教学之类的方式)</span></div><div class="notion-text notion-block-39737eaec0b181c7b02defe76504fe87"><span class="notion-default">自蒸馏没什么好说的，用下面的公式计算loss，然后更新就行了，借鉴MESA的方法。本文同时还用了EMA更新teacher</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b1819cb97ee25950bc2b51"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://pic3.zhimg.com/v2-8b0bd4a9bc5271b3ff32854b2c54177a_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-819c-b97e-e25950bc2b51" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181fa94bec08ce57e1e28"><span class="notion-default">这里主要想提的是这个优化器，前面优化器为SGD，后面阶段是AdamW，看回答下面的大佬有提到，这个优化方案在SGD中会比较有效，AdamW效果貌似不佳。其更新方法如下所示，来自大佬在</span><span class="notion-default"><a class="notion-link" href="https://www.zhihu.com/question/638766873/answer/3358801861" target="_blank" rel="noopener noreferrer">深度学习中，loss下降的快慢或者曲率（但最后收敛在同一水平）会对下游任务的性能有什么影响吗？</a></span><span class="notion-default">回答下面的code</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181fcb2a3d1c738863182" data-id="39737eaec0b181fcb2a3d1c738863182"><span><div id="39737eaec0b181fcb2a3d1c738863182" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181fcb2a3d1c738863182" title="总结："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">总结：</span></span></span></h3><div class="notion-text notion-block-39737eaec0b18157bdf9e3f73bd40723"><span class="notion-default">论文中的不少方式现在都已经属于训练的通用设置了，我就不怎么提了，刚好因为自己现在在做low-level，训练集和测试集只能说毫不相关，正好看见这个优化器，有种耳目一新的感觉，下午来尝试一下效果。我个人还是比较喜欢这种，简单有效的工作。相比许多理论一大堆，复现困难不说，还很难拓展应用到自己工作的论文，好上太多了</span></div><div class="notion-text notion-block-39737eaec0b181869513cb7d547b179b"><span class="notion-default">尝试了一下ALRS算法，由于我是在low-level上使用的，会遇到一个问题就是SGD基本不收敛，可以看到蓝色的是使用了SGDM优化的模型，而其他是正常AdamW的模型，loss上有着巨大的差距。无奈放弃。</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b181828cbdc91e9ac60ee8"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://picx.zhimg.com/v2-637df68378d3f1e5231c770313846a1e_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8182-8cbd-c91e9ac60ee8" alt="notion image" loading="lazy" decoding="async"/></div></figure></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[光流][ECCV2024 Oral]SEA-RAFT: Simple, Efficient, Accurate RAFT  for Optical Flow]]></title>
            <link>https://www.zeurd.com/article/zhihu-1969851071414375676</link>
            <guid>https://www.zeurd.com/article/zhihu-1969851071414375676</guid>
            <pubDate>Thu, 06 Nov 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[最近在看ICCV oral的时候，发现了一篇关于光流的paper，准备研读一下，就发现这篇SEA-RAFT是其前置base model，仔细一看还有点意思，并且其方案复现简单，就自己尝试了一下，确实有效，就想着写一篇文章简单介绍一下。 论文地址： https://arxiv.org/abs/2405.14793Moti…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b181c69817eb0ba11a9ed9"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181ab90d0dea2398a3f1c"/><div class="notion-text notion-block-39737eaec0b1816691fffc2711c31415"><span class="notion-default">最近在看ICCV oral的时候，发现了一篇关于光流的paper，准备研读一下，就发现这篇SEA-RAFT是其前置base model，仔细一看还有点意思，并且其方案复现简单，就自己尝试了一下，确实有效，就想着写一篇文章简单介绍一下。</span></div><div class="notion-text notion-block-39737eaec0b1817ba8e3d3b843a74602"><span class="notion-default">论文地址：</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1812890bfd724700d931b" data-id="39737eaec0b1812890bfd724700d931b"><span><div id="39737eaec0b1812890bfd724700d931b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1812890bfd724700d931b" title="Motivation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Motivation</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181adab66e1a67d0f7590"><span class="notion-default">作者提出了一种简单有效且准确的RAFT光流新方法。（没啥motivation，就是数值碾压）也符合CV会议的精髓，对于故事会的要求比ACL低多了。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181be8778e6a79f7cca60" data-id="39737eaec0b181be8778e6a79f7cca60"><span><div id="39737eaec0b181be8778e6a79f7cca60" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181be8778e6a79f7cca60" title="Contribution"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Contribution</span></span></span></h3><ul class="notion-list notion-list-disc notion-block-39737eaec0b181b89203eeadf6432a5d"><li><span class="notion-default">作者提出了一个Mixture of Laplace Loss</span></li></ul><ul class="notion-list notion-list-disc notion-block-39737eaec0b1812182fbca2901ca5777"><li><span class="notion-default">作者不是从0开始初始化，而是预测一个初始流来初始化</span></li></ul><ul class="notion-list notion-list-disc notion-block-39737eaec0b1812eb30be09addc7249c"><li><span class="notion-default">使用了刚性光流数据集做预训练</span></li></ul><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1815d95edc848c912fb29" data-id="39737eaec0b1815d95edc848c912fb29"><span><div id="39737eaec0b1815d95edc848c912fb29" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1815d95edc848c912fb29" title="Method"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default"><b>Method</b></span></span></span></h3><div class="notion-text notion-block-39737eaec0b18175b3adcbef502a25eb"><span class="notion-default"><b>Mixture of Laplace Loss：</b></span></div><div class="notion-text notion-block-39737eaec0b181f3b41af71b0f852601"><span class="notion-default">这个loss其实就是这篇论文性能提升的主要来源。loss的假设是，光流本身具有不确定性，导致很多时候存在模棱两可的光流，有时候位移一个像素是对的，位移两个像素也是对的，这时候如果强制使用L1去学习的话，可能会学会一个偏执的，过拟合于训练集的错误光流，导致结果在测试集上不鲁棒。</span></div><div class="notion-text notion-block-39737eaec0b18173baf8f6150163db7b"><span class="notion-default">作者认为这种时候的分布是基于拉普拉斯分布的，于是作者在预测光流的2个通道的同时，还预测了4个通道，他们分别是b通道以及weight通道。其中b通道是拉普拉斯分布的参数b，weight是一个区分的通道。作者的解释是用单一的拉普拉斯分布效果不好，所以他们用了2个拉普拉斯。</span></div><div class="notion-text notion-block-39737eaec0b1814e9281d34087ba04aa"><span class="notion-default">打个比方来说，你在给同学批改作业，每道题都是一个像素的答案，和标准答案有误差。大部分题目清楚易判，错一点就该扣分，所以让严格老师的来评分，错一点就狠狠的罚。而对于困难的题目，本身就模棱两可，很难打死板的分，那就让宽容的老师来评分。但是我们也不知道到底哪个题目是清楚的，哪个是困难的，所以让模型自己学了一个，来使用softmax来决定让哪个老师评分，这就是weight通道。简单来说，weight通道决定了方向（严格还是宽容），而b通道决定了振幅。（有多严厉以及有多宽容）为了让数值稳定一点，作者还用了对数形式来约束b。</span></div><div class="notion-text notion-block-39737eaec0b1811798ced69e067d535f"><span class="notion-default">公式可以直接看论文，说白了就是两个拉普拉斯分布mask加权求和，没啥意思</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b181d79381eae5baf2c262"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://pic3.zhimg.com/v2-90c38d68a2fbf3b53382e97f47372854_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-81d7-9381-eae5baf2c262" alt="notion image" loading="lazy" decoding="async"/></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18182ad8cf834f6fc1efc"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://pica.zhimg.com/v2-294a50ccecdac9c4eacca783c85bf2f6_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8182-ad8c-f834f6fc1efc" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181039e90e558190683a1"><span class="notion-default">这个loss的具体实现源码里有现成的，不过在模型的forward里，不大常规，当时看的时候找了有好一会。地址在这儿</span><span class="notion-default"><a class="notion-link" href="https://github.com/princeton-vl/SEA-RAFT/blob/main/core/raft.py" target="_blank" rel="noopener noreferrer">https://github.com/princeton-vl/SEA-RAFT/blob/main/core/raft.py</a></span><span class="notion-default"> 感兴趣可以自己看一下，实现起来挺简单的。</span></div><div class="notion-text notion-block-39737eaec0b1817cb104c1c4969af2a2"><span class="notion-default"><b>Direct Regression of Initial Flow：</b></span></div><div class="notion-text notion-block-39737eaec0b181758212e959bcd56e14"><span class="notion-default">其实介绍完了这个loss，这篇论文已经没啥东西了，这里有一个初始流回归，说白了就是原来要初始化一个全0的张量然后在上面做refine，现在直接不用初始化了，前面的网络直接预测一个初始化的光流，然后在上面refine，感觉上除了加快收敛速度，其他没啥效果，事实上我自己做光流一直都是直接预测的，没想到这也能成为一个contribution。</span></div><div class="notion-text notion-block-39737eaec0b1813a9fb3da928971d8d3"><span class="notion-default"><b>Large-Scale Rigid-Flow Pre-Training：</b></span></div><div class="notion-text notion-block-39737eaec0b1810194c7d11308eb1ba9"><span class="notion-default">使用刚性数据集，这个contirbution就更奇怪了，说实话我有点get不到，在我理解里，就是用了一个比flychair更简单的数据集做预训练，然后逐步换数据集加难度。</span></div><div class="notion-text notion-block-39737eaec0b181abb2d9d66284fcf042"><span class="notion-default">光流任务经常有这种，先训A，再训B，然后训C这种训练方式，看了这么多论文其实我一直没get到到底和直接一起训有啥区别，反正我自己先xxx，再xxx，和两个一起训久一点，在自己的实验里没发现有什么区别过。可能是属于肉眼感知不到，结果数值有优化的trick吧</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181cebc67f6df8064577c" data-id="39737eaec0b181cebc67f6df8064577c"><span><div id="39737eaec0b181cebc67f6df8064577c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181cebc67f6df8064577c" title="Conclusion"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Conclusion</span></span></span></h3><div class="notion-text notion-block-39737eaec0b1813189d8df13bb821827"><span class="notion-default">这篇论文就写完了，当时看了挺久，还上手实现了一下。确实属于简单有效的方法，尤其是4通道在推理的时候还能丢掉。不增加额外算力，属于是0代价提升性能了。</span></div><div class="notion-text notion-block-39737eaec0b18158a6a6e52ad3a39eea"><span class="notion-default">不过这篇论文其实也存在一些问题的，比如：这个loss说白了，就是加权算loss，只不过是逐像素加权罢了。那么让模棱两可的部分权重小了，如果这个模棱两可的地方正好很重要应该怎么办？毕竟工程不是写论文，真用起来可不是给甲方报个数字就行了。又比如，这个方法刚需光流gt，但是事实上大部分光流任务都是自监督无监督的，有监督gt的光流数据集真的不多，而且性能受限于gt光流也是一个问题。还有一个就是，raft的通病，开销比较大，端侧基本用不了。（今年ICCV要处理的问题，下篇讲）</span></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[CVPR 2025 Best Paper]VGGT: Visual Geometry Grounded Transformer]]></title>
            <link>https://www.zeurd.com/article/zhihu-1927029499431744954</link>
            <guid>https://www.zeurd.com/article/zhihu-1927029499431744954</guid>
            <pubDate>Tue, 31 Mar 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[最近在看图像深度相关的内容，就准备刷刷论文看一看，随便一搜，就看见这篇 VGGT: Visual Geometry Grounded Transformer 拿了去年CVPR的best paper，吓了我一跳还以为是VGG又崛起了，又要make CNN great again了，赶忙下下来看了一眼，就浅浅略读一下。本…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b1816baeb5c6b27f2346c8"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b18107bca4f91ae4add55c"/><div class="notion-text notion-block-39737eaec0b18118ab3feb928e66200d"><span class="notion-default">最近在看图像深度相关的内容，就准备刷刷论文看一看，随便一搜，就看见这篇</span><span class="notion-default"><a class="notion-link" href="https://cvpr.thecvf.com/virtual/2025/poster/33969" target="_blank" rel="noopener noreferrer">VGGT: Visual Geometry Grounded Transformer</a></span><span class="notion-default"> 拿了去年CVPR的best paper，吓了我一跳还以为是VGG又崛起了，又要make CNN great again了，赶忙下下来看了一眼，就浅浅略读一下。</span></div><div class="notion-text notion-block-39737eaec0b1810db0c1f87dcf941710"><span class="notion-default">本篇主要想用比较直白的话来讲述一下VGGT干了啥，怎么干，目标是让和我一样，没接触过这个方向的同学也能大致了解一下到底在干啥。有些细节可能了解的不是很到位，欢迎专门做3D重建的大佬指正。</span></div><div class="notion-text notion-block-39737eaec0b1811eb08ae23b9aa6a052"><span class="notion-default">顺带说一句，看完论文感觉真就突出一个力大砖飞，我还特意看了眼单位，meta的，看的时候我还以为是谷歌的呢。meta给llama4不行找到借口了，搁这儿训VGGT呢。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181718554c167f132f719" data-id="39737eaec0b181718554c167f132f719"><span><div id="39737eaec0b181718554c167f132f719" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181718554c167f132f719" title="Motivation"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Motivation</span></span></span></h3><div class="notion-text notion-block-39737eaec0b1810f8c7cc3a24194cdd3"><span class="notion-default">VGGT的motivation非常简单易懂，就是传统三维重建太复杂了，极其依赖几何计算，即使是NN的方法也依旧需要大量的几何处理。所以VGGT的motivation非常直接，能不能用网络直接完成三维重建，避免复杂的几何运算。然后以此开始研究怎么端到端三维重建。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18167b308c3f799ea5568" data-id="39737eaec0b18167b308c3f799ea5568"><span><div id="39737eaec0b18167b308c3f799ea5568" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18167b308c3f799ea5568" title="Method"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Method</span></span></span></h3><div class="notion-text notion-block-39737eaec0b1819b88d3c5ea6c2ef563"><span class="notion-default">VGGT的method在我这个外行看起来，其实相当简单，不知道是不是有什么在三维重建领域非常标新立异的方法，有懂的朋友也可以指出来。</span></div><div class="notion-text notion-block-39737eaec0b181a8a828e1fa9c549689"><span class="notion-default"><b>Feature Backbone</b></span></div><div class="notion-text notion-block-39737eaec0b1818ab88ddc5e00fb5adf"><span class="notion-default">简单来说， 就是使用了一个VIT的架构，使用 DINO 将每帧切成 patch-tokens，再把所有帧 token 串起来输入VIT网络。
同时，略微修改了一下VIT，引入了Alternating-Attention的机制，一层 Frame-wise SA，一层 Global SA，循环往复 24 次。</span></div><div class="notion-text notion-block-39737eaec0b181a2a0e1d580f2b145c6"><span class="notion-default"><b>Prediction heads</b></span></div><div class="notion-text notion-block-39737eaec0b181c5b829e4210efc0b36"><span class="notion-default">网络的预测那更是直接一网打尽，把相机参数、深度图、点云图 及点轨迹一次全部出来。具体操作上，有趣的是每个token 队列都加了1个 Camera Token ＋ 4 个 Register Tokens，在输出的时候再把register token丢弃。而且只有第一帧用了一套可学习的Token，同时其他帧用另一套，显示的告诉模型从哪里开始。</span></div><div class="notion-text notion-block-39737eaec0b181798091ed75ed6f89d2"><span class="notion-default">然后就是各个种类的一预测头了，用4个注意力层最后做个先行曾得到相机参数预测，用DPT得到深度，点云图。用CoTracker2架构得到追踪特征。这个地方我感觉应该属于即插即用，随时可以替换成最新的sota任务的头然后fintune，以保持一个有竞争力的效果。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1815594d6d6103249f1ef" data-id="39737eaec0b1815594d6d6103249f1ef"><span><div id="39737eaec0b1815594d6d6103249f1ef" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1815594d6d6103249f1ef" title="实验："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">实验：</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181f199dbf5b144fe3741"><span class="notion-default">实验主要做了大量的实验，对自己每个头对应的任务进行了实验，结果么，一言以蔽之，赢麻了。</span></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[量化]HAWQ]]></title>
            <link>https://www.zeurd.com/article/zhihu-719063685</link>
            <guid>https://www.zeurd.com/article/zhihu-719063685</guid>
            <pubDate>Mon, 09 Sep 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[本文讲述了HAWQ系列的三部曲，HAWQ v1，v2，v3 论文名称： HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision (ICCV 2019) HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks (NeurIPS 2020) HAWQ-V3: Dyad…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b18190a0cfdf22fd4e9ded"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b18169828bdcc818321671"/><div class="notion-text notion-block-39737eaec0b181929505c878d9a27bc6"><span class="notion-default">本文讲述了HAWQ系列的三部曲，HAWQ v1，v2，v3</span></div><div class="notion-text notion-block-39737eaec0b1810cb1cfd0c7c6e383b7"><span class="notion-default">论文名称：</span><span class="notion-default"><a class="notion-link" href="https://openaccess.thecvf.com/content_ICCV_2019/html/Dong_HAWQ_Hessian_AWare_Quantization_of_Neural_Networks_With_Mixed-Precision_ICCV_2019_paper.html" target="_blank" rel="noopener noreferrer"><span class="notion-inline-underscore">HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision (ICCV 2019)</span></a></span></div><div class="notion-text notion-block-39737eaec0b181bd92d4e9223bff7e2e"><span class="notion-default"><a class="notion-link" href="https://proceedings.neurips.cc//paper/2020/file/d77c703536718b95308130ff2e5cf9ee-Paper.pdf" target="_blank" rel="noopener noreferrer"><span class="notion-inline-underscore">HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks (NeurIPS 2020)</span></a></span></div><div class="notion-text notion-block-39737eaec0b181cabfa5c55ffd456c96"><span class="notion-default"><a class="notion-link" href="https://arxiv.org/abs/2011.10680" target="_blank" rel="noopener noreferrer"><span class="notion-inline-underscore">HAWQ-V3: Dyadic Neural Network Quantization (ICML 2021)</span></a></span></div><div class="notion-text notion-block-39737eaec0b181e98562d65bf768e07d"><span class="notion-default"><span class="notion-inline-underscore">项目地址：</span></span><span class="notion-default"><a class="notion-link" href="https://github.com/Zhen-Dong/HAWQ" target="_blank" rel="noopener noreferrer">https://github.com/Zhen-Dong/HAWQ</a></span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1817e8b23ebe458d90c84" data-id="39737eaec0b1817e8b23ebe458d90c84"><span><div id="39737eaec0b1817e8b23ebe458d90c84" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1817e8b23ebe458d90c84" title="故事会环节"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">故事会环节</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181f1ad46ee8d565b4eb7"><span class="notion-default">大背景说一下就是模型越来越大，特别是输入分辨率越来越大，所以降低模型的大小提高速度非常重要，然后写了一堆现在大家怎么做的，其中一种方案就是量化。也就是本文的重点。然后指出来，所有的网络都量化到一个bit会降精度，所以提出混合精度量化，对敏感层用高精度，非敏感层用低精度。所以在V1当中，提出了整个系列的关键做法，使用海森矩阵，计算Hessian谱，确定量化水平。在V2当中，作者就V1中的一些不足继续补充。</span></div><div class="notion-text notion-block-39737eaec0b181b9af4bf347888f744d"><span class="notion-default">V1存在几个缺点，（i） HAWQ 使用基于顶部 Hessian 特征值的启发式度量作为灵敏度的度量，它忽略了 Hessian 谱的其余部分;（ii） HAWQ 仅提供不同层的相对灵敏度，仍然需要手动选择混合精度设置;（iii） HAWQ 不考虑混合精度激活量化。为了解决这些问题，作者提出了V2，解决以上3个问题：（i）使用平均Hessian轨迹，而不是启发式算法。（ii）使用基于 Pareto-frontier 的方法自动确定不同层的位精度。（iii）开发了混合精度激活量化，提出了快速计算Hssian信息的方法，计算激活。</span></div><div class="notion-text notion-block-39737eaec0b18105966fcc01a2af8a1c"><span class="notion-default">V2呢，当然也有几个缺点，（1）一个是虽然量化了，但是在推理过程中  的操作，还是一个浮点计算，在硬件上还是慢.（2）V2的Pareto-frontier在不同硬件上表现有差异。所以又提出了V3，解决这两个问题，（1）使用了移位操作来代替除法运算，使得整个网络是纯整数推理。（2）提出了使用一个整数规划来通过求方程得到到底解决到底使用什么精度配置的问题。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181d4a87ef648017ff696" data-id="39737eaec0b181d4a87ef648017ff696"><span><div id="39737eaec0b181d4a87ef648017ff696" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181d4a87ef648017ff696" title="创新点:"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">创新点:</span></span></span></h3><div class="notion-text notion-block-39737eaec0b18159b5b0cda5f22c3d5e"><span class="notion-default">总结三篇论文，最终能用的创新点其实是如下几个，其中有一些创新点已经被迭代掉了。</span></div><ol start="1" class="notion-list notion-list-numbered notion-block-39737eaec0b1815fbdbdeb6df8e8fe80" style="list-style-type:decimal"><li><span class="notion-default">使用平均Hessian轨迹来计算层的灵敏度。（V2）</span></li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39737eaec0b181ba9d9ec491a1889d0c" style="list-style-type:decimal"><li><span class="notion-default">使用了移位操作来代替除法运算，使得整个网络是纯整数推理（V3）</span></li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-39737eaec0b181b38326cc31d6afc0f6" style="list-style-type:decimal"><li><span class="notion-default">使用整数规划来确定到底哪些层4bit，哪些层8bit</span></li></ol><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18111936bee5fd0da589b" data-id="39737eaec0b18111936bee5fd0da589b"><span><div id="39737eaec0b18111936bee5fd0da589b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18111936bee5fd0da589b" title="Method"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Method</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181eabb7dd61ac32a3fad"><span class="notion-default">对于已经被迭代掉的方法暂且放下不提，本文只考虑最后被应用的方法。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181d7903ad2de9a2639fb" data-id="39737eaec0b181d7903ad2de9a2639fb"><span><div id="39737eaec0b181d7903ad2de9a2639fb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181d7903ad2de9a2639fb" title="平均Hessian轨迹"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">平均Hessian轨迹</span></span></span></h4><div class="notion-text notion-block-39737eaec0b1815082d3cd135b1bce34"><span class="notion-default">这个证明非常长，作者写了2页纸来证明这个公式，有兴趣的小伙伴可以自行去论文阅读。虽然读了貌似也不大明白，建议直接看代码，比推理看着好理解多了。</span></div><div class="notion-text notion-block-39737eaec0b181068df9c6f3928a8acd"><span class="notion-default">经过一个很长很长的公式推理之后，论文得出了平均Hessian轨迹</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b1813fa324dc4f631117ad"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img src="https://pica.zhimg.com/v2-1834a41a031bb4667692925da3c93634_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-813f-a324-dc4f631117ad" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181ab8bd6e33fd5184a9d"><span class="notion-default">近似一下，估算的公式可以被看做是</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18125bedff27dc4b68b72"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic3.zhimg.com/v2-46cdafcd5dcf63f33d7283a87b0fe74a_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8125-bedf-f27dc4b68b72" alt="notion image" loading="lazy" decoding="async"/></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b18198abfcd10866caf90c" data-id="39737eaec0b18198abfcd10866caf90c"><span><div id="39737eaec0b18198abfcd10866caf90c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18198abfcd10866caf90c" title="纯整数推理"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">纯整数推理</span></span></span></h4><div class="notion-text notion-block-39737eaec0b181f5ad43cba8d21d34c5"><span class="notion-default">简单来说，就是把一个  的值，用  毕竟，然后作为scale，优点是真的快，因为没有除法而且是纯整数。缺点是会进一步掉精度。</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18193b4d9ee599a7b581d"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic3.zhimg.com/v2-72df22f75894fa514ef68aec401bb1ab_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8193-b4d9-ee599a7b581d" alt="notion image" loading="lazy" decoding="async"/></div></figure><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b18181be52fc6120a705bb" data-id="39737eaec0b18181be52fc6120a705bb"><span><div id="39737eaec0b18181be52fc6120a705bb" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18181be52fc6120a705bb" title="量化层选择"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">量化层选择</span></span></span></h4><div class="notion-text notion-block-39737eaec0b181b8a83fd37764e740b1"><span class="notion-default">这一点简单来说，就是把一个模型部署在一个设备上，必然有权重大小，推理延迟，带宽限制。说白了就是前面用Hessian矩阵计算了灵敏度，得到了一个层从重要到不重要的排序，但是中间的临界点在哪里，那就说不准了。所以作者说，很科学，要使用一个整数规划，去计算不同层的权重大小，推理延迟，带宽限制。然后如下所示，通过解一个整数规划来计算到底哪些层用4bit，哪些用8bit</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18194bc4ce13aa2e0cace"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic1.zhimg.com/v2-531bd461b4a073bb3e771632298703ea_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8194-bc4c-e13aa2e0cace" alt="notion image" loading="lazy" decoding="async"/></div></figure><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181a78989fd867eb81019" data-id="39737eaec0b181a78989fd867eb81019"><span><div id="39737eaec0b181a78989fd867eb81019" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181a78989fd867eb81019" title="总结："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">总结：</span></span></span></h3><div class="notion-text notion-block-39737eaec0b18112abb9ee67a86012b1"><span class="notion-default">作者的核心观点在于一个假设，就是通过Hessian矩阵轨迹，可以得到一个层的灵敏度，从而得出哪些层可以用低精度，哪些需要高精度。之后的一系列改进，都是在这个基础上，怎么更快，怎么更准。于是设计了纯整数量化，做整数规划等等。本文最大的价值其实在我看来反而是最后一个，考虑到了工程部署的问题，学界做的真的能落地能应用的东西真的太少了。</span></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[多任务学习][NIPS2020]Gradient Surgery for Multi-Task Learning]]></title>
            <link>https://www.zeurd.com/article/zhihu-14457677240</link>
            <guid>https://www.zeurd.com/article/zhihu-14457677240</guid>
            <pubDate>Tue, 24 Dec 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[最近在跑模型的时候，遇到了个问题，单独训练A loss ，模型可以达到想要的性能，训练B loss，也可以达到想要的性能，但是当A和B一起的时候，就两个任务都达不成了。搜了搜这个方向的论文，发现了这篇一千都引用的paper，试了一下，确实有点作用，故而记录一…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b18118b63fdc166a4db169"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181a6b46dcc10516bced1"/><div class="notion-text notion-block-39737eaec0b181ebb3d2ffab62778637"><span class="notion-default">最近在跑模型的时候，遇到了个问题，单独训练A loss ，模型可以达到想要的性能，训练B loss，也可以达到想要的性能，但是当A和B一起的时候，就两个任务都达不成了。搜了搜这个方向的论文，发现了这篇一千都引用的paper，试了一下，确实有点作用，故而记录一下。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18182bc82e105db2cd53a" data-id="39737eaec0b18182bc82e105db2cd53a"><span><div id="39737eaec0b18182bc82e105db2cd53a" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18182bc82e105db2cd53a" title="论文题目：Gradient Surgery for Multi-Task Learning"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">论文题目：</span><span class="notion-default"><a class="notion-link" href="https://proceedings.neurips.cc/paper_files/paper/2020/hash/3fe78a8acf5fda99de95303940a2420c-Abstract.html" target="_blank" rel="noopener noreferrer">Gradient Surgery for Multi-Task Learning</a></span></span></span></h3><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18163904fecc59a9662f0" data-id="39737eaec0b18163904fecc59a9662f0"><span><div id="39737eaec0b18163904fecc59a9662f0" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18163904fecc59a9662f0" title="代码地址：https://github.com/tianheyu927/PCGrad"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">代码地址：</span><span class="notion-default"><a class="notion-link" href="https://github.com/tianheyu927/PCGrad" target="_blank" rel="noopener noreferrer">https://github.com/tianheyu927/PCGrad</a></span></span></span></h3><div class="notion-text notion-block-39737eaec0b181bd907ce33c507c17b2"><span class="notion-default">因为是用于训练的论文，这里就不说很多有的没的故事了。简单来说，就是当发现在跑多任务的时候，如果发下任务冲突（文中的判断方式是余弦相似度为负）。那么就不能正常训练了。于是作者想了个办法，就是在传梯度的时候，对梯度做点小手术，把负向的梯度向下图所示，投影到正向上，这样不就余弦相似度是正的了么。然后试了一下，效果确实还挺管用。</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b181e29d79e4ee1a696a73"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pica.zhimg.com/v2-94588a3e4e21713b242452d5b027000b_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-81e2-9d79-e4ee1a696a73" alt="notion image" loading="lazy" decoding="async"/></div></figure><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18185b2b5f2d522763bdf"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic4.zhimg.com/v2-579e3857ce08d6325e56af8fea7a1328_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8185-b2b5-f2d522763bdf" alt="公式1" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption"><span class="notion-default">公式1</span></figcaption></div></figure><div class="notion-text notion-block-39737eaec0b181998199ff838ce1d2ed"><span class="notion-default">论文的方法核心就是用公式1来解决梯度冲突的问题。不过官网的代码是tf的版本，我自己手写了一个pytorch的，方便使用。</span></div><div class="notion-text notion-block-39737eaec0b1815c9936c4bbc2e425d9"><span class="notion-default">上面代码是根据tf的代码化简写的，和本体可能有一定出入，仅供参考。</span></div><div class="notion-text notion-block-39737eaec0b1817abe0dfba0bd007228"><span class="notion-default">不过在使用pcgrad的时候，也是遇到了一些问题，特别的有下面2个点</span></div><ul class="notion-list notion-list-disc notion-block-39737eaec0b181849496dc61b956b2f7"><li><span class="notion-default">对于多任务，作者代码中使用的是随机选取主任务然后对其他做梯度手术，但是一般做多个loss的时候，其实是有主导性的，哪个loss比较重要，哪个没这么重要，都有说法，直接随机可能导致最重要的任务训练不充分。</span></li></ul><ul class="notion-list notion-list-disc notion-block-39737eaec0b181898cb2e3ed0e4a7d28"><li><span class="notion-default">在做梯度手术的时候，对于内积为正，余弦方向相同的loss作者选择的是直接舍弃。但是就算方向完全相同，也不代表另一个loss不需要了，如果不加就能达成目的那根本就不需要多出来这额外的loss。这可能导致有些loss一直得不到优化。</span></li></ul><div class="notion-text notion-block-39737eaec0b181c98cfae0ca48cf2591"><span class="notion-default">后续准备再看看改进，有没有更好的解决方案吧，不过这篇论文工作确实不错，方法简单好用，值得推荐</span></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[如何从0开始，学习AI（二）]]></title>
            <link>https://www.zeurd.com/article/zhihu-2426588400</link>
            <guid>https://www.zeurd.com/article/zhihu-2426588400</guid>
            <pubDate>Tue, 22 Oct 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[最近忙着做项目，一直没时间更新，今天抽出点时间写一下第二节。 在上一节 如何从0开始，学习AI（一）中，我大概阐述了一下应该如何入门机器学习，深度学习算法，以及如何入门AI的代码，并讲述了到什么程度，就叫入门了。在本节中，我会讲述应该如何从0开始…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-39737eaec0b181409e48c413e19eebad"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181d58c29c32f7284f162"/><div class="notion-text notion-block-39737eaec0b181ad8de9fa1f953e4d96"><span class="notion-default">最近忙着做项目，一直没时间更新，今天抽出点时间写一下第二节。</span></div><div class="notion-text notion-block-39737eaec0b1811ba71eed30db1f1fd3"><span class="notion-default">在上一节 </span><span class="notion-default"><a class="notion-link" href="https://zhuanlan.zhihu.com/p/721810749" target="_blank" rel="noopener noreferrer">如何从0开始，学习AI（一）</a></span><span class="notion-default">中，我大概阐述了一下应该如何入门机器学习，深度学习算法，以及如何入门AI的代码，并讲述了到什么程度，就叫入门了。在本节中，我会讲述应该如何从0开始，去阅读一篇论文，以及到底应该阅读哪些论文。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1810bb614e4ba4c88cced" data-id="39737eaec0b1810bb614e4ba4c88cced"><span><div id="39737eaec0b1810bb614e4ba4c88cced" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1810bb614e4ba4c88cced" title="如何选择阅读的论文"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">如何选择阅读的论文</span></span></span></h3><div class="notion-text notion-block-39737eaec0b1817988d3dd1eb01d9906"><span class="notion-default">在人工智能领域，论文数量浩如烟海，每天arxiv上更新的论文，就算一个人不吃不喝不睡，也不可能读的完。所以我们在阅读论文的时候，必然需要进行有选择性的阅读，才能在有限的时间里，吸收到足够的信息。</span></div><div class="notion-text notion-block-39737eaec0b1816faa77fa352e448fef"><span class="notion-default">而关于论文阅读的目标，一般来说其实分为2个方向，一个是</span><span class="notion-default"><b>找灵感</b></span><span class="notion-default">，一个是</span><span class="notion-default"><b>学写作</b></span><span class="notion-default">。</span></div><div class="notion-text notion-block-39737eaec0b181a993d6f213da0c7765"><span class="notion-default">找灵感说白了，就是找idea，如果没有老师给你安排idea的话，基本idea就是自己看论文找的了。而且这其实也是日后，不管是自己做科研，还是从事算法工作，最重要的能力。网上能看见很多博士论文发不出来，说白了就是硕士的时候，老师帮助太大，可能自己就读了几篇论文，老师直接丢给他一个课题他照着做，做完把结果写成论文，然后老师改完就中了所谓的顶会，申请读了博士。但是其实他自己压根没有独立发现问题，解决问题，找到灵感的能力，在读了博士之后，甚至教职之后，需要自己做科研的时候，就发现脑袋空空。</span></div><div class="notion-text notion-block-39737eaec0b181dfb45ec9a2db4852c5"><span class="notion-default">另一个就是学习写作，事实上，一篇论文能不能中，很大程度取决于你的写作水平，我自己在当时刚开始写作的时候，投的NIPS，然后在review阶段就被审稿人痛骂，说是虽然想法很新颖，实验结果很solid，但是这篇论文的写作让他怀疑实验结果的真实性，然后就低分out了。说白了，因为AI论文本身的验证，特别是现在LLM，验证成本以及难度很高，所以审稿人在审稿的时候，判断你的论文真实性最重要的依据，一个是你的idea的逻辑链条听起来是不是符合审稿人的认知，另一个则是你的写作水平，反应出的科学素养。</span></div><div class="notion-text notion-block-39737eaec0b18145b528d09f77e9eb55"><span class="notion-default">那么到底要如何选择，能帮助找灵感，提升写作能力的论文呢？</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181619d3ed1be14888fc5" data-id="39737eaec0b181619d3ed1be14888fc5"><span><div id="39737eaec0b181619d3ed1be14888fc5" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181619d3ed1be14888fc5" title="找灵感"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">找灵感</span></span></span></h4><div class="notion-text notion-block-39737eaec0b1810e823ec6eb77eff4cb"><span class="notion-default">对于找灵感的论文，总共分为两步</span></div><ol start="1" class="notion-list notion-list-numbered notion-block-39737eaec0b1811996b6e13e22810aaa" style="list-style-type:decimal"><li><span class="notion-default">找到准备做的方向</span></li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39737eaec0b18197bf62ef8f599a3498" style="list-style-type:decimal"><li><span class="notion-default">大量阅读相关顶会的工作</span></li></ol><div class="notion-text notion-block-39737eaec0b1815b92a8c4ceb18f1305"><span class="notion-default">第一步，也是最重要的，要找到一个自己准备从事的方向，组里有传承，就在组里传承的基础上接着做，如果组里没传承，需要自己找方向的，那就根据自己能得到的计算资源来选择研究方向。有传承的不谈，我只能留下羡慕的泪水，本文主要针对的是没有传承的同学。对于没有传承的同学来说，最重要的就是估算自己能调度的计算资源，然后选一个稍微热门一点的方向来做。</span></div><div class="notion-text notion-block-39737eaec0b181a982e4e04655e5611d"><span class="notion-default">计算资源的原理很简单，你只有一块4090，甚至3090的卡，你妄想去做LLM相关的工作，这不是老寿星吃砒霜么？你的计算资源直接决定了你能做啥，至于怎么判断，每一篇论文的实验设置部分，都会写论文的显卡配置，找自己能承受的来。</span></div><div class="notion-text notion-block-39737eaec0b18180a755ddba1b572fc9"><span class="notion-default">至于选热门的方向，这个就仁者见仁了，AI领域有很多占山头的情况，那种冷门的小方向，很有可能整个方向都被某个组包圆了，只有他们自己在做，审稿的时候大家一看标题可能就大概知道是不是哪个组做的了，那你一个圈外人想发，难度就很大了。但是热门方向，就要考虑有没有可能做的人很多，你的一个idea有可能有很多组在和你一起做，最后就是分秒必争了。</span></div><div class="notion-text notion-block-39737eaec0b181c0932cd18903339179"><span class="notion-default">典型的就是只要一个新的backbone出来，那关于它的，下游迁移，各种常规轻量化改造，马上就会在一两个月，两三个月内被人写了发arxiv抢山头。你出了xxx，明天就已经有人新建latex在准备mobilexxx，shufflexxx，resxxx，xxx for seg，xxx for detection，就都出来了。虽然没啥营养，但是确实能发呀，而且这种没啥营养的论文，万一这个backbone后来火了，他们也能分一杯羹，引用就高了。这就和打个王者，出一个新英雄就开始一堆主播冲国服，一般都不是技术主播，但是没事，新的英雄国服分低，比的是速度不是实力，最后这英雄出久了，他们小国都进不去，但是不要紧，历史标记是国标了已经。</span></div><div class="notion-text notion-block-39737eaec0b181ef97ace7080401e5eb"><span class="notion-default">在找到了方向之后，那就要开始大量的阅读顶会里的相关工作了。一开始建议读relate-work，一般如果是负责的论文，会把到底这篇论文为啥这么做，是从哪些工作找的灵感，都写的很清楚。当然，也有不负责的，比如中科大前两天不报出来说ICLR被评价为抄袭desk reject，部分原因就是因为relate-work直接照抄的，非常偷懒。评判这种其实也很简单，你看一下它的写作脉络，有没有那种娓娓道来的感觉，还是无脑的xxx用了xxx，xxx用了xxx，其实还挺明显的，多看几篇就知道了。了解整个脉络之后，就开始读这些论文的intro，看看他们领域里有什么问题，他们找到了什么问题，怎么解决的。</span></div><div class="notion-text notion-block-39737eaec0b181c3a5cff070d230ef39"><span class="notion-default">如果准备做小创新，那就沿着别人的思路做，看论文解决了问题之后，还有哪里不足，然后再想想怎么改。如果做中创新，那就是这个领域有什么问题，他这么做解决了，但是这么做不行，自己提出一个解决方法。如果要做大创新，那大哥应该不需要看我的入门，毕竟我自己好像也没做出过啥创新来。</span></div><div class="notion-text notion-block-39737eaec0b181238ce1e0773de20246"><span class="notion-default">在找到灵感之后，那就需要开始读提升写作能力的论文了。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b18181b759efb66896b433" data-id="39737eaec0b18181b759efb66896b433"><span><div id="39737eaec0b18181b759efb66896b433" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18181b759efb66896b433" title="学写作"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">学写作</span></span></span></h4><div class="notion-text notion-block-39737eaec0b181dbb372dae010a26836"><span class="notion-default">很多朋友其实会很疑惑，学写作和找灵感，都是读论文，有区别么？为什么要分开说？这其实就涉及到一个准备发啥的问题。</span></div><div class="notion-text notion-block-39737eaec0b181bb8f78ccc14166c826"><span class="notion-default">因为很多同学的学校，是不认会议的，CVPR=EI会议＜SCI4区的事情在我们AI领域的学校里时有发生，你费劲巴拉好不容易发了一篇会议，结果发现评奖学金不认，毕业不认，那不尬住了？所以学习写作，最重要的还是要找到目标期刊。你要投会议，里面也差别很大，比如ICLR今年是10页，AAAI就只有6页还是7页，ICASSP更是只有4页，而期刊更是普遍不限制页数，只要你付得起版面费，那你写多少都行。版面能放多少东西，直接决定了你要写多少，论文每一个部分的轻重缓急，都不一样，需要对着你的目标会和目标期刊，从那里面找论文，然后学习写作，才有实际意义。更有甚者，我还遇到过实验室有人，投了个3区想水一篇，结果审了好几个月被打回来，因为写太好审稿人让转投更好的期刊的，结果一审又是几个月，浪费了大量时间。最后在毕业答辩前看看收到录取通知，当时他都已经写好中文核心准备投中文保毕业了，可想而知选错的代价有多大。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181bb8c9ee43a50def44f" data-id="39737eaec0b181bb8c9ee43a50def44f"><span><div id="39737eaec0b181bb8c9ee43a50def44f" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181bb8c9ee43a50def44f" title="如何选择阅读的论文"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">如何选择阅读的论文</span></span></span></h3><div class="notion-text notion-block-39737eaec0b18162a586c9a694eeb5fa"><span class="notion-default">在选择好自己要读的论文后，自然就要开始具体的读论文，这里又分成了几个部分，分别是关于每一块的阅读的侧重点。abs暂且不谈，主要是为了让你了解这篇论文是干啥的，有没有点进去的必要，属于必须通篇阅读的东西。其他的大致分为Introduction，Relate-work，Method，Experiment几个部分，其他有各种写法的，大致都可以归结到这里面去。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b1814ba876c53115eedeb1" data-id="39737eaec0b1814ba876c53115eedeb1"><span><div id="39737eaec0b1814ba876c53115eedeb1" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1814ba876c53115eedeb1" title="Introduction"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Introduction</span></span></span></h4><div class="notion-text notion-block-39737eaec0b1815f9eddfa9d6f613e76"><span class="notion-default">Introduction可以说是一篇论文最重要的部分，一篇好论文的Introduction，可以非常清晰的勾勒出作者在思考这篇论文的时候的想法。论文的背景是什么？在这个背景下有什么问题？前人是怎么解决问题的？还有哪些不足?作者提出了什么解决了这些不足？大抵上是这样的一个流程。</span></div><div class="notion-text notion-block-39737eaec0b1815eaf67f6c7b79ddc23"><span class="notion-default">阅读一篇introduction，最关键的部分，也就是上述问题之间的逻辑衔接了，背景问题不足这些东西阐述的是否有一个完整的逻辑链条，自己的方法是否真的解决了自己提出的问题，这些都是Introduciton需要考虑的，我们阅读一篇Introduction也正是如此。而判断一篇工作到底是不是有意义足够优秀，其实也主要看的是这些。别人在找论文的时候，第一个看到就是你提出的现在有的问题是不是别人遇到的问题，如果正好是，那你工作的意义就出来了。其次就是看有没有有效，如果正好有效被别人采用了，那就是一篇论文最大的贡献。但是如果一篇论文，它的大背景接的问题和背景没关系，问题也不是这个领域真的存在的问题，那大概率就是一篇不怎么样的论文。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181b18331e13ca288cc3c" data-id="39737eaec0b181b18331e13ca288cc3c"><span><div id="39737eaec0b181b18331e13ca288cc3c" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181b18331e13ca288cc3c" title="Relate-work"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Relate-work</span></span></span></h4><div class="notion-text notion-block-39737eaec0b181868dfeccec5b6551fd"><span class="notion-default">Relate-work主要是阐述一下作者是通过哪些工作得到的灵感，顺便给不是很了解这个领域的人了解一下大致情况，如果已经很熟悉了，我自己感觉可以不看。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b1812a897ad311ee3dbe16" data-id="39737eaec0b1812a897ad311ee3dbe16"><span><div id="39737eaec0b1812a897ad311ee3dbe16" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1812a897ad311ee3dbe16" title="Method"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Method</span></span></span></h4><div class="notion-text notion-block-39737eaec0b18174923ffbddc78d7f6f"><span class="notion-default">论文的method部分，一般是在论文精读的时候看，我个人不是很赞成每篇论文的Method都看，因为这其实是一篇论文里看起来最费时的事，展现了论文具体是怎么实现的。就我个人习惯来说，我会现在Introduction，如果我觉得这篇论文是对我有帮助的，我才会看一下论文的Method，看下具体是怎么做的。同时还得看看abs有没有提供代码，如果没有代码，method又写的很抽象的，我也一样不看，因为我知道可能复现起来比我自己想一个写还累，同时还不一定真的有用，但是如果有代码，对照着代码看method，理解起来会方便很多。（当然，那种屎山代码，看了半天连作者Idea写在哪儿都不知道的除外）</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181c788e7d65c28c17aba" data-id="39737eaec0b181c788e7d65c28c17aba"><span><div id="39737eaec0b181c788e7d65c28c17aba" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181c788e7d65c28c17aba" title="Experiment"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">Experiment</span></span></span></h4><div class="notion-text notion-block-39737eaec0b181ee81bde511c2ea7f1b"><span class="notion-default">实验部分我问过好几个朋友他们是怎么看的，问下来的结果大差不大，直接看图看表就行了，文字部分其实没有那么重要，大部分都是对表格内容的阐述。</span></div><div class="notion-text notion-block-39737eaec0b181f28929d83b36abcf8e"><span class="notion-default">但是一篇论文看到Experiment，必须要有的一个概念就是，你需要对论文的实验结果有一个大概的预期，然后看实验部分来印证一下自己的预期，或者说，让你自己来设计实验，你会怎么设计？这比如我在Introduction的时候看论文说，做的轻量化工作，或者降低了内存开销，那就会着重看实验部分里这些内容的实验，有没有符合我心里的预期，或者结果怪不怪。前几天就遇到了一篇做降低内存开销的paper，结果实验通篇在测准确率，我就有点无语，是因为和刷点的比准确率比不过，所以换个赛道比么？但是人家做轻量化的也没一定要准确率呀，估计是两头都不拔尖，所以选择不比了。这种工作大概率就不是很硬。但是如果你对论文在读到这儿没有一个自己的预期的话，那就发现每篇论文结果都是最好的，实验看不看都一样，毕竟不好的结果他也不可能放出来。</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b181bbba92fba227f117c9" data-id="39737eaec0b181bbba92fba227f117c9"><span><div id="39737eaec0b181bbba92fba227f117c9" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181bbba92fba227f117c9" title="小结"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">小结</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181419e0cc21813970a8b"><span class="notion-default">读论文的方法大概总结到这儿，如果能从浩如烟海的论文中找灵感，学写作，然后对一篇论文能有自己的理解与解释，我个人认为阅读论文的能力就已经非常可以了。下一期，看后续忙不忙吧，可能会出如何写作，也可能就懒更了，继续记录自己看的论文，看心情哈哈。确实感觉上班不比上学的时候松弛了，可惜没念个博，以后看机会能不能念一个哈哈哈</span></div></main></div>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[[降噪][ECCV 2020]Practical Deep Raw Image Denoising on Mobile Devices]]></title>
            <link>https://www.zeurd.com/article/zhihu-716143319</link>
            <guid>https://www.zeurd.com/article/zhihu-716143319</guid>
            <pubDate>Fri, 23 Aug 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[论文题目： Practical Deep Raw Image Denoising on Mobile Devices故事会环节：传统的去噪网络很难在移动设备上部署，所以设计了一个能在一动设备上部署使用的轻量化去噪模型。（过去模型呢，需要在不同噪声水平训练不同的模型使用，不然盲降噪性能很差）…]]></description>
            <content:encoded><![CDATA[<div id="notion-article" class="mx-auto overflow-hidden "><main class="notion light-mode notion-page notion-block-9d205b087ff547329ac71ba47bb3242c"><div class="notion-viewport"></div><div class="notion-collection-page-properties"></div><hr class="notion-hr notion-block-39737eaec0b181a19b17c788d538b417"/><div class="notion-text notion-block-39737eaec0b1817ca2a3d3285234861d"><span class="notion-default">论文题目：</span><span class="notion-default"><a class="notion-link" href="https://link.springer.com/chapter/10.1007/978-3-030-58539-6_1" target="_blank" rel="noopener noreferrer">Practical Deep Raw Image Denoising on Mobile Devices</a></span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b1817b9e63e9a378528f31" data-id="39737eaec0b1817b9e63e9a378528f31"><span><div id="39737eaec0b1817b9e63e9a378528f31" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b1817b9e63e9a378528f31" title="故事会环节："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">故事会环节：</span></span></span></h3><div class="notion-text notion-block-39737eaec0b181e6a323e57f902b0fcc"><span class="notion-default">传统的去噪网络很难在移动设备上部署，所以设计了一个能在一动设备上部署使用的轻量化去噪模型。（过去模型呢，需要在不同噪声水平训练不同的模型使用，不然盲降噪性能很差），本文使用了k-sigma变换将不同的ISO变换到一个空间。性能达到了sota</span></div><h3 class="notion-h notion-h2 notion-h-indent-0 notion-block-39737eaec0b18187b62ad9d0276350bf" data-id="39737eaec0b18187b62ad9d0276350bf"><span><div id="39737eaec0b18187b62ad9d0276350bf" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18187b62ad9d0276350bf" title="创新点"><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">创新点</span></span></span></h3><ol start="1" class="notion-list notion-list-numbered notion-block-39737eaec0b18153bca9cf4e2767fafc" style="list-style-type:decimal"><li><span class="notion-default">这是旷视出的用于商业的模型，确实能用（能用在水文漫天的AI领域里也算重中之重了）</span></li></ol><ol start="2" class="notion-list notion-list-numbered notion-block-39737eaec0b181e9b6d2d7ccd0a2ff95" style="list-style-type:decimal"><li><span class="notion-default">提出了一个k-sigma变换，用来把不同水平的噪声归一化到同一个域里面，用一个模型解决不同的噪声问题</span></li></ol><ol start="3" class="notion-list notion-list-numbered notion-block-39737eaec0b181a0a50cdc558ce3d26d" style="list-style-type:decimal"><li><span class="notion-default">提出一个轻量化的网络</span></li></ol><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181b6b6ecd8e6c1bf3855" data-id="39737eaec0b181b6b6ecd8e6c1bf3855"><span><div id="39737eaec0b181b6b6ecd8e6c1bf3855" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181b6b6ecd8e6c1bf3855" title="k-sigma变换："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">k-sigma变换：</span></span></span></h4><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b18199a6c7dacde9a4c9e9"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://picx.zhimg.com/v2-43622ed9a96b6d2a05ebbc7b12af704c_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-8199-a6c7-dacde9a4c9e9" alt="notion image" loading="lazy" decoding="async"/></div></figure><div class="notion-text notion-block-39737eaec0b181b28b8afb3f84972acc"><span class="notion-default">这里对于具体的证明过程就不细说了，文章中写的很细，细到从光子成像开始说了都，可以好好的去读一下。简单来说就是作者通过对高斯泊松噪声的分解，分析，标定，简化，得到了一个结论，</span><em><span class="notion-default">k与sigma只和g相关，g</span></em><span class="notion-default">是由</span><em><span class="notion-default">ISO</span></em><span class="notion-default">决定的，通过</span><em><span class="notion-default">k-sigma</span></em><span class="notion-default">变换，可以把噪声建模成一个和</span><em><span class="notion-default">ISO</span></em><span class="notion-default">无关的表示，从而把不同等级的噪声放到一个网络里去训，做到盲降噪。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b181a88e2af69bb0347231" data-id="39737eaec0b181a88e2af69bb0347231"><span><div id="39737eaec0b181a88e2af69bb0347231" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b181a88e2af69bb0347231" title="网络结构："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">网络结构：</span></span></span></h4><div class="notion-text notion-block-39737eaec0b18191a038eb7680fcdfbf"><span class="notion-default">网络具体结构如下：</span></div><figure class="notion-asset-wrapper notion-asset-wrapper-image notion-block-39737eaec0b181d2a2e8fd6986830684"><div style="position:relative;display:flex;justify-content:center;align-self:center;width:100%;max-width:100%;flex-direction:column"><img style="object-fit:cover" src="https://pic3.zhimg.com/v2-ddcd994dc224a61617d8e8bfeb0c46c4_r.jpg?spaceId=1ae623ef-62da-428f-a135-e32993b17333&amp;t=39737eae-c0b1-81d2-a2e8-fd6986830684" alt="a图为整体网络结构，b图为一些块的具体细节" loading="lazy" decoding="async"/><figcaption class="notion-asset-caption"><span class="notion-default">a图为整体网络结构，b图为一些块的具体细节</span></figcaption></div></figure><div class="notion-text notion-block-39737eaec0b181fc85daf11e7967cb37"><span class="notion-default">简单来说，就是一个经典的U-net架构（虽然基本上像素级任务都是这么个架构）特点是加上了一个Sep block，用了个5×5的卷积核增加了点感受野。在长连接的时候也进行了一些操作，没有直接加上来，但整体大差不差把，把身子缩小把支路放大的操作。</span></div><h4 class="notion-h notion-h3 notion-h-indent-1 notion-block-39737eaec0b18157a301e4e4f62af05b" data-id="39737eaec0b18157a301e4e4f62af05b"><span><div id="39737eaec0b18157a301e4e4f62af05b" class="notion-header-anchor"></div><a class="notion-hash-link" href="#39737eaec0b18157a301e4e4f62af05b" title="总结："><svg viewBox="0 0 16 16" width="16" height="16"><path fill-rule="evenodd" d="M7.775 3.275a.75.75 0 001.06 1.06l1.25-1.25a2 2 0 112.83 2.83l-2.5 2.5a2 2 0 01-2.83 0 .75.75 0 00-1.06 1.06 3.5 3.5 0 004.95 0l2.5-2.5a3.5 3.5 0 00-4.95-4.95l-1.25 1.25zm-4.69 9.64a2 2 0 010-2.83l2.5-2.5a2 2 0 012.83 0 .75.75 0 001.06-1.06 3.5 3.5 0 00-4.95 0l-2.5 2.5a3.5 3.5 0 004.95 4.95l1.25-1.25a.75.75 0 00-1.06-1.06l-1.25 1.25a2 2 0 01-2.83 0z"></path></svg></a><span class="notion-h-title"><span class="notion-default">总结：</span></span></span></h4><div class="notion-text notion-block-39737eaec0b1819e9233f931ae3f95b7"><span class="notion-default">总结来说，这篇论文最大的创新点就是提出了一个k-sigma变换，把不同噪声放到同一个域里，至于网络结构，只能说U-net NB吧，不过k-sigma变换是基于高斯泊松噪声的前提推导的，如果噪声标定不是用的这种方式，或者直接就是真实噪声，不确定效果</span></div></main></div>]]></content:encoded>
        </item>
    </channel>
</rss>