Rectified Flow VQVAE
[x] Vanilla RF VQVAE 성능 측정
→ Tradeoff between stochastic / deterministic sampling in FID / MPJPE
→ Effect of reflow is not significant
→ Implemented with very basic MLP (BxSxD → B*SxD) : No Attention Used
[x] Architecture Improvements
[ ] UNet / Transformer (or at least some attention in temporal dimension)
[ ] End to End Training (Following Sample what you cannot compress scheme)
[x] Text Conditioning (추후에는 반드시 필요할 것이라고 생각됨)
[x] Reflow 의 유효성 검증
[ ] 다양한 파라미터, 디자인 검증
Conditioning 보다는 좋다라는게 논문에서 보여주기는 함. (그리고 굳이 Pretrained World Model이 없는데 Conditioning 할 필요는 없어보인다.)
→ 이거에 대한 설명이 있긴했네 (Condition 시키면 Posterior MSE랑 동치다)
[ ] Theoretical Background
Vanilla (2D)VQVAE + RF 에 대한 T2M 성능 측정
[ ] MMM 스킴을 그대로 따라서 Bidirectional 1 Stage Generation (일단 1D로 Flatten해서 빠르게 확인?) (Fixed Length) → WIP
[ ] 결국 우리가 2DVQVAE를 가지고 있기도 하고 성능이 좋다는게 충분히 보여진것 같으므로 MogenTS 의 2D Token Map + 2D Masking Strategy Handling (단 밑에 Autoregressive + Bidirectional 세팅과의 연결도 고려해야함.)
[ ] BAMM의 Autoregressive + Bidirectional 세팅 차용
[ ] Architecture 고정하면 VQVAE 디자인 등 바꿔가면서 Ablation 해보기
[ ] 왜 VQVAE가 잘되는가? 에 대한 고찰 (Like Cross Entropy)
→ VQVAE + Continuous refinement 라는 Pipeline 을 우리가 왜 사용해야하는가?
Diffusion VQVAE
MAR