본문으로 건너뛰기
jongkwan.dev

Posts

2개 포스트
  1. Multi-head Latent Attention과 KV 캐시 압축

  2. Transformer 구조와 LLM 서빙 최적화