AI & Algorithms
Attention Is All You Need:Transformer 架构完整拆解
拆解原始 Transformer 论文的 scaled dot-product attention、multi-head attention、encoder/decoder 区块、残差与 Layer Normalization、位置编码与训练方式,并区分 2017 年原始论文与现代 decoder-only LLM 在架构细节上的实际差异。
阅读全文深入 AI 与数学、系统与编程、网络与基础设施、数据与存储。
浏览所有文章AI & Algorithms
拆解原始 Transformer 论文的 scaled dot-product attention、multi-head attention、encoder/decoder 区块、残差与 Layer Normalization、位置编码与训练方式,并区分 2017 年原始论文与现代 decoder-only LLM 在架构细节上的实际差异。
阅读全文AI & Algorithms
从随机变量、随机向量和样本路径出发,建立随机过程与随机场的基本直觉。
Networking & Infrastructure
记录 sing-box VLESS+Reality 透明代理迁移到 ASUSWRT-Merlin 路由器时的架构、TPROXY、IPv6、回滚、启动时序与 SQM 经验。
按最后更新时间
AI & Algorithms
Systems & Programming
Networking & Infrastructure
Networking & Infrastructure
Systems & Programming
Systems & Programming
Networking & Infrastructure