CUDA SGEMM 并行优化 cuda · gpu · performance · systems 从朴素矩阵乘法出发,用共享内存分块优化 SGEMM,并通过多尺寸基准测试分析吞吐、冲突和异常点。 报告记录 512×512 测试达到 999.51 GFLOPS,相对 baseline 提升 85.5%。 cuda gpu performance systems
昇腾 910C 训练与推理兼容性诊断 llm-infrastructure · systems-engineering · ascend · performance · compatibility 在团队大模型工程中参与训练性能排查、benchmark 选型与算子兼容修复,并严格区分个人贡献和团队结果。 个人完成三类输入格式的动态路由修复;团队记录显示 MFU 从约 2% 提升至 25%-30%。 llm-infrastructure systems-engineering ascend performance compatibility