Vision transformer
id:
vision-transformer-204-15898543
title:
Vision transformer
text:
A vision transformer (ViT) is a transformer designed for computer vision. A ViT breaks down an input image into a series of patches, serialises each patch into a vector, and maps it to a smaller dimension with a single matrix multiplication. These vector embeddings are then processed by a transformer encoder as if they were token embeddings. ViT were designed as alternatives to convolutional neural networks (CNN) in computer vision applications. They have different inductive biases, training sta
brand slug:
wiki
category slug:
encyclopedia
description:
Variant of Transformer designed for vision processing
original url:
https://en.wikipedia.org/wiki/Vision_transformer
date created:
2021-07-11T18:52:09Z
date modified:
2024-09-10T05:23:36Z
main entity:
{"identifier":"Q107675654","url":"https://www.wikidata.org/entity/Q107675654"}
image:
{"content_url":"https://upload.wikimedia.org/wikipedia/commons/9/93/Vision_Transformer.png","width":1426,"height":788}
fields total:
13
integrity:
16