Skip to content

[5] X-ViT: High Performance Linear Vision Transformer without Softmax #5

Description

@techzzt

X-ViT: High Performance Linear Vision Transformer without Softmax
image
Figure 3. X-ViT module

  • Computer vision task에서 기존의 self-attention (SA) algorithm의 complexity를 최소화하며 학습하는 ViT 구조를 제안
  • 본 논문에서 제안하는 알고리즘인 X-ViT는 기존의 SA에 대해 nonlinearity를 제거한 모델
  • 기존의 ViT 코드에서 몇 줄만 변경했음에도 ImageNet Top-1 accuracy에 대해 Swin 모델, DeiT 모델에 비해 향상된 결과를 보임

image
Figure 1. Top-1 accuracy vs. Model capacity

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions