<?xml version='1.0' encoding='UTF-8'?>
<doi_batch version="5.4.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="http://www.crossref.org/schema/5.4.0" xsi:schemaLocation="http://www.crossref.org/schema/5.4.0 https://www.crossref.org/schemas/crossref5.4.0.xsd" xmlns:jats="http://www.ncbi.nlm.nih.gov/JATS1" xmlns:fr="http://www.crossref.org/fundref.xsd" xmlns:ai="http://www.crossref.org/AccessIndicators.xsd" xmlns:rel="http://www.crossref.org/relations.xsd" xmlns:mml="http://www.w3.org/1998/Math/MathML">
  <head>
    <doi_batch_id>NONE</doi_batch_id>
    <timestamp>20260610115421857</timestamp>
    <depositor>
      <depositor_name>wseas/wseas</depositor_name>
      <email_address>content-registration-form+ja@crossref.org</email_address>
    </depositor>
    <registrant>content-registration-form</registrant>
  </head>
  <body>
    <journal>
      <journal_metadata>
        <full_title>International Journal of Computational and Applied Mathematics &amp; Computer Science</full_title>
        <issn media_type="electronic">2769-2477</issn>
      </journal_metadata>
      <journal_article>
        <titles>
          <title>Big Data Analytics in Healthcare: Tools, Challenges and Opportunities</title>
        </titles>
        <contributors>
          <person_name sequence="first" contributor_role="author">
            <given_name>Evangelia N.</given_name>
            <surname>Petraki</surname>
            <affiliations>
              <institution>
                <institution_name>Department of Economics National and Kapodistrian University of Athens Sofokleous 1 Str, 105 59 GREECE </institution_name>
              </institution>
            </affiliations>
          </person_name>
        </contributors>
        <jats:abstract>
          <jats:p>The term big data has been used in recent years to describe the enormous volume of data collected by various means and at a very rapid pace in all areas of human activity. Big data differs from traditional data sets and requires special management due to its volume, structure and type. For this reason, specialized technologies for their processing are being developed. The use of machine learning methods to take advantage of the big data collected in the health sector is imperative as it can significantly contribute to making predictions, preventing illnesses, improving patient healthcare, reducing costs and optimizing the use of available resources. This paper deals with Big Data, briefly presenting its characteristics, the sources of its creation, as well as the basic technologies used for its processing and utilization in the health sector. In the second part of this work, medical data on diabetes were used, different software tools were integrated to get information using SQL queries and machine learning methods were applied using the Weka open-source software to evaluate and compare the results for diabetes prediction. The paper concludes with useful concepts arising from the analysis, as well as challenges and suggestions for future research.</jats:p>
        </jats:abstract>
        <publication_date media_type="print">
          <month>06</month>
          <day>10</day>
          <year>2026</year>
        </publication_date>
        <publication_date media_type="online">
          <month>06</month>
          <day>10</day>
          <year>2026</year>
        </publication_date>
        <pages>
          <first_page>11</first_page>
        </pages>
        <publisher_item>
          <item_number item_number_type="article_number">2</item_number>
        </publisher_item>
        <ai:program name="AccessIndicators">
          <ai:license_ref>https://creativecommons.org/licenses/by/4.0/deed.en_US</ai:license_ref>
        </ai:program>
        <doi_data>
          <doi>10.37394/232028.2026.6.2</doi>
          <resource>https://wseas.com/journals/camcs/2026/a04camcs-002(2026).pdf</resource>
        </doi_data>
        <citation_list>
          <citation key="ref0">
            <unstructured_citation>H. Basim Alwan, K.R. Ku-Mahamud, “Big data: definition, characteristics, life cycle, applications, and challenges IOP”, Conference Series: Materials Science and Engineering, 769 (1) (2020), doi:10.1088/1757-899X/769/1/012007</unstructured_citation>
          </citation>
          <citation key="ref1">
            <unstructured_citation>A. Gandomi, M. Haider, “Beyond the hype: Big data concepts, methods, and analytics”, International Journal of Information Management, Volume 35, Issue 2, 2015, Pages 137-144, ISSN 0268-4012, https://doi.org/10.1016/j.ijinfomgt.2014.10.0 07</unstructured_citation>
          </citation>
          <citation key="ref2">
            <unstructured_citation>B.Furht, F. Villanustre, “Big Data Technologies and Applications”, 2016, Springer International Publishing, ISBN 978- 3-319-44548-9, DOI 10.1007/978-3-319- 44550-2.</unstructured_citation>
          </citation>
          <citation key="ref3">
            <unstructured_citation>D. Gupta, R. Rani, “A study of big data evolution and research challenges”, 2019, Journal of Information Science, 45(3), 322– 340. https://doi.org/10.1177/0165551518789880</unstructured_citation>
          </citation>
          <citation key="ref4">
            <unstructured_citation>Adriana Alexandru, Cristina Adriana Alexandru, Dora Coardos, Eleonora Tudora, "Healthcare, Big Data and Cloud Computing," WSEAS Transactions on Computer Research, vol. 4, pp. 123-131, 2016.</unstructured_citation>
          </citation>
          <citation key="ref5">
            <unstructured_citation>Belle A., Thiagarajan R., Soroushmehr SM, Navidi F., Beard DA, Najarian K (2015), «Big data analytics in healthcare». Biomed Research International, Volume 2015, Article ID 370194, http://dx.doi.org/10.1155/2015/370194</unstructured_citation>
          </citation>
          <citation key="ref6">
            <unstructured_citation>Raghupathi W, Raghupathi V., (2014) “Big data analytics in healthcare: promise and potential”. Health Information Science and Systems, 7;2:3. doi: 10.1186/2047-2501-2-3. PMID: 25825667; PMCID: PMC4341817.</unstructured_citation>
          </citation>
          <citation key="ref7">
            <unstructured_citation>Guo, C., Chen, J. (2023). Big Data Analytics in Healthcare. In: Nakamori, Y. (eds) Knowledge Technology and Systems. Translational Systems Sciences, vol 34. Springer, Singapore. https://doi.org/10.1007/978-981-99-1075- 5_2</unstructured_citation>
          </citation>
          <citation key="ref8">
            <unstructured_citation>Batko, K., Ślęzak, A. (2022) The use of Big Data Analytics in healthcare. J Big Data, 9, 3. https://doi.org/10.1186/s40537-021-00553-4</unstructured_citation>
          </citation>
          <citation key="ref9">
            <unstructured_citation>Geerts, G. L., &amp; O'Leary, D. E. (2022). V-Matrix: A wave theory of value creation for big data. International Journal of Accounting Information Systems, Elsevier, vol. 47(C).</unstructured_citation>
          </citation>
          <citation key="ref10">
            <unstructured_citation>Pouchard, L. (2015). Revisiting the data lifecycle with big data curation. International Journal of Digital Curation 10 (2), 176-192.</unstructured_citation>
          </citation>
          <citation key="ref11">
            <unstructured_citation>Jagadish, H. V., Gehrke, J., Labrinidis, A., Papakonstantinou, Y., Patel, J. M., Ramakrishnan, R., &amp; Shahabi, C. (2014). Big data and its technical challenges. Communications of the ACM, 57(7), 86-94.</unstructured_citation>
          </citation>
          <citation key="ref12">
            <unstructured_citation>Apache storm, https://storm.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref13">
            <unstructured_citation>Apache Flume, https://flume.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref14">
            <unstructured_citation>Apache Kafka, https://kafka.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref15">
            <unstructured_citation>OpenRefine, https://openrefine.org/ (accessed 14 of September 2025)</unstructured_citation>
          </citation>
          <citation key="ref16">
            <unstructured_citation>MongoDB, https://www.mongodb.com/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref17">
            <unstructured_citation>Apache Cassandra, https://cassandra.apache.org/_/index.html (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref18">
            <unstructured_citation>Apache Spark, https://spark.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref19">
            <unstructured_citation>Tableau, https://www.tableau.com/products/desktop (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref20">
            <unstructured_citation>Power BI https://www.microsoft.com/enus/power-platform/products/power-bi (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref21">
            <unstructured_citation>Python, https://www.python.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref22">
            <unstructured_citation>Weka, https://ml.cms.waikato.ac.nz/weka/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref23">
            <unstructured_citation>Apache Hadoop Distributed FileSystem (HDFS) https://hadoop.apache.org/docs/r1.2.1/hdfs_d esign.html (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref24">
            <unstructured_citation>Apache Hadoop Mapreduce https://hadoop.apache.org/docs/r1.2.1/mapre d_tutorial.html (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref25">
            <unstructured_citation>Hive, https://hive.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref26">
            <unstructured_citation>Kevin S. Beyer, Vuk Ercegovac, Rainer Gemulla, Andrey Balmin, Mohamed Eltabakh, Carl-Christian Kanne, Fatma Ozcan, and Eugene J. Shekita. 2011. Jaql: a scripting language for large scale semistructured data analysis. Proc. VLDB Endow. 4, 12 (08/2011), 1272–1283. https://doi.org/10.14778/3402755.3402761</unstructured_citation>
          </citation>
          <citation key="ref27">
            <unstructured_citation>Zookeeper, https://zookeeper.apache.org/ (accessed 14 of September 2025).</unstructured_citation>
          </citation>
          <citation key="ref28">
            <unstructured_citation>Diabetes Health Indicators Dataset https://www.kaggle.com/datasets/alexteboul/ diabetes-health-indicators-dataset File diabetes _ 012 _ health _ indicators _ BRFSS2015.csv, a clean dataset of 253,680 survey responses to the CDC's BRFSS2015. (accessed 31 of May 2026).</unstructured_citation>
          </citation>
          <citation key="ref29">
            <unstructured_citation>Eclipse Temurin jdk-17.0.19 https://adoptium.net/temurin/releases?versio n=17&amp;os=any&amp;arch=any (accessed 31 of May 2026).</unstructured_citation>
          </citation>
          <citation key="ref30">
            <unstructured_citation>Diabetes Dataset Published: 18 July 2020| Version 1 | DOI: 10.17632/wj9rwkp9c2.1 Contributor: Ahlam Rashid https://data.mendeley.com/datasets/wj9rwkp 9c2/1 (accessed 31 of May 2026).</unstructured_citation>
          </citation>
        </citation_list>
      </journal_article>
    </journal>
  </body>
</doi_batch>
