<?xml version="1.0" encoding="UTF-8"?>
<doi_batch xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.crossref.org/schema/5.5.0 https://www.crossref.org/schemas/crossref5.5.0.xsd" xmlns="http://www.crossref.org/schema/5.5.0" xmlns:jats="http://www.ncbi.nlm.nih.gov/JATS1" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:ai="http://www.crossref.org/AccessIndicators.xsd" version="5.5.0">
  <head>
    <doi_batch_id>wseas-12080-20260923082850-c5efb7</doi_batch_id>
    <timestamp>20260923082850911</timestamp>
    <depositor>
      <depositor_name>wseas/wseas</depositor_name>
      <email_address>wseas.group@gmail.com</email_address>
    </depositor>
    <registrant>WSEAS</registrant>
  </head>
  <body>
    <journal>
      <journal_metadata language="en">
        <full_title>WSEAS Transactions on Computers</full_title>
        <issn media_type="print">1109-2750</issn>
        <issn media_type="electronic">2224-2872</issn>
      </journal_metadata>
      <journal_issue>
        <publication_date media_type="online">
          <month>04</month>
          <day>14</day>
          <year>2026</year>
        </publication_date>
        <publication_date media_type="print">
          <month>04</month>
          <day>14</day>
          <year>2026</year>
        </publication_date>
        <journal_volume>
          <volume>25</volume>
        </journal_volume>
      </journal_issue>
      <journal_article publication_type="full_text" language="en">
        <titles>
          <title>Understanding Vision Transformers through Intuition and Simple Visual Examples</title>
        </titles>
        <contributors>
          <person_name sequence="first" contributor_role="author">
            <given_name>M.</given_name>
            <surname>Sabrigiriraj</surname>
            <affiliations>
              <institution>
                <institution_name>Department of It Hindusthan College of Engineering and Technology Coimbatore INDIA</institution_name>
              </institution>
            </affiliations>
          </person_name>
          <person_name sequence="additional" contributor_role="author">
            <given_name>K.</given_name>
            <surname>Manoharan</surname>
            <affiliations>
              <institution>
                <institution_name>Department of Ece Sns College of Technology Coimbatore INDIA</institution_name>
              </institution>
            </affiliations>
          </person_name>
        </contributors>
        <jats:abstract xml:lang="en"><jats:p>Vision Transformers are normally used in image recognition and are considered as an alternative to Convolutional Neural Networks. Unlike traditional convolution-based models, a Vision Transformer does not rely on local filters Instead, it segments an image into smaller patches and understands the relationships among individual patches. This technique is powerful, but it can be challenging for beginners to understand when described only with equations or programming concepts. Hence, this paper showcases the process of Vision Transformers using simple examples and visual illustrations. This paper is intended for students, teachers, and beginners who are interested in learning how transformer models are used in computer vision. The tutorial explains how an image is divided into patches, how each patch is transformed into a token, how self-attention helps different image regions share information, and how the transformer encoder helps the final classification process. The foremost purpose of this tutorial is to support readers in developing a clear understanding of Vision Transformers. After gaining this intuitive understanding, readers can more easily study the mathematical particulars and implementation techniques later.</jats:p></jats:abstract>
        <publication_date media_type="online">
          <month>09</month>
          <day>23</day>
          <year>2026</year>
        </publication_date>
        <publication_date media_type="print">
          <month>09</month>
          <day>23</day>
          <year>2026</year>
        </publication_date>
        <pages>
          <first_page>160</first_page>
        </pages>
        <publisher_item>
          <item_number item_number_type="article_number">15</item_number>
        </publisher_item>
        <ai:program name="AccessIndicators">
          <ai:free_to_read/>
          <ai:license_ref applies_to="vor" start_date="2026-09-23">https://creativecommons.org/licenses/by/4.0/</ai:license_ref>
        </ai:program>
        <doi_data>
          <doi>10.37394/23205.2026.25.15</doi>
          <resource>https://wseas.com/journals/articles.php?id=12080</resource>
        </doi_data>
        <citation_list>
          <citation type="journal_article" key="ref1"><unstructured_citation>M. Sabrigiriraj, K. Manoharan. Teaching Machine Learning and Deep Learning Introduction: An Innovative Tutorial-Based Practical Approach. WSEAS Transactions on Advances in Engineering Education. Vol. 21, pp. 54–61, 2024.10.37394/232010.2024.21.8</unstructured_citation></citation>
          <citation type="journal_article" key="ref2"><unstructured_citation>S. Khan, M. Naseer, M. Hayat, S. Waqas Zamir, F. S. Khan and M. Shah &quot;Transformers in vision: A survey.&quot; ACM computing surveys (CSUR) Vol. 54, no. 10s, pp. 1–41, 2022.</unstructured_citation></citation>
          <citation type="journal_article" key="ref3"><unstructured_citation>K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, E. Xu, C. Xu, Y. Xu and D. Tao &quot;A survey on vision transformer.&quot; IEEE transactions on pattern analysis and machine intelligence Vol. 45, no. 1, pp. 87–110, 2023.</unstructured_citation></citation>
          <citation type="journal_article" key="ref4"><unstructured_citation>Y. Liu, T. Zhang, K. Chen, X. Liu, Y. Deng and F. Wang, &quot;A survey of visual transformers.&quot; IEEE transactions on neural networks and learning systems Vol. 35, no. 6, pp. 7478–7498, 2024.</unstructured_citation></citation>
          <citation type="journal_article" key="ref5"><unstructured_citation>Jamil, Sonain, Md Jalil Piran, and Oh-Jin Kwon. &quot;A comprehensive survey of transformers for computer vision.&quot; drones, Vol. 7, no. 5, Art. no. 287, 2023.</unstructured_citation></citation>
          <citation type="journal_article" key="ref6"><unstructured_citation>A. Khan, M. Khan, S. Khan and M. A. Khan &quot;A survey of the self-supervised learning mechanisms for vision transformers.&quot; arXiv preprint arXiv:2408. 17059 (2024).</unstructured_citation></citation>
          <citation type="journal_article" key="ref7"><unstructured_citation>Fournier, Quentin, Gaétan Marceau Caron, and Daniel Aloise. &quot;A practical survey on faster and lighter transformers.&quot; ACM Computing Surveys Vol. 55, no. 14s, pp. 1–40, 2023.</unstructured_citation></citation>
          <citation type="journal_article" key="ref8"><unstructured_citation>R. Kashefi, M. Shafiee, A. M. N. Nik and M. H. Mahoor, &quot;Explainability of vision transformers: A comprehensive review and new perspectives.&quot; arXiv preprint arXiv:2311.06786 (2023).</unstructured_citation></citation>
          <citation type="journal_article" key="ref9"><unstructured_citation>Khalil, Mahmoud, Ahmad Khalil, and Alioune Ngom. &quot;A comprehensive study of vision transformers in image classification tasks.&quot; arXiv preprint arXiv:2312.01232 (2023).</unstructured_citation></citation>
          <citation type="journal_article" key="ref10"><unstructured_citation>K. Al-Hammuri, M. Alhaj, M. Alsmirat and A. Al-Ayyoub, &quot;Vision transformer architecture and applications in digital health: a tutorial and survey.&quot; Visual computing for industry, biomedicine, and art , Vol. 6, no. 1, Art. no. 14, 2023.</unstructured_citation></citation>
          <citation type="journal_article" key="ref11"><unstructured_citation>Y. Yang, Y. Wu, M. Wang, H. Hu, L. Lin and B. Du, &quot;Transformers meet visual learning understanding: A comprehensive review.&quot; arXiv preprint arXiv:2203.12944 (2022).</unstructured_citation></citation>
          <citation type="journal_article" key="ref12"><unstructured_citation>Maurício, José, Inês Domingues, and Jorge Bernardino. &quot;Comparing vision transformers and convolutional neural networks for image classification: A literature review.&quot; Applied Sciences Vol. 13, no. 9, Art. no. 5521, 2023.</unstructured_citation></citation>
          <citation type="journal_article" key="ref13"><unstructured_citation>L. Papa, G. Fiameni, L. Rundo, C. Distante and I. M. Ray, &quot;A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking.&quot; IEEE transactions on pattern analysis and machine intelligence Vol. 46, no. 12, pp. 7682–7700, 2024.</unstructured_citation></citation>
        </citation_list>
      </journal_article>
    </journal>
  </body>
</doi_batch>
