<?xml version="1.0" encoding="UTF-8"?>
<doi_batch version="5.3.1" xmlns="http://www.crossref.org/schema/5.3.1" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:jats="http://www.ncbi.nlm.nih.gov/JATS1" xmlns:ai="http://www.crossref.org/AccessIndicators.xsd" xsi:schemaLocation="http://www.crossref.org/schema/5.3.1 http://www.crossref.org/schema/deposit/crossref5.3.1.xsd">
 <head>
  <doi_batch_id>aspg-3-730-1791417024</doi_batch_id>
  <timestamp>20261007235024</timestamp>
  <depositor>
   <depositor_name>American Scientific Publishing Group</depositor_name>
   <email_address>admin@americaspg.com</email_address>
  </depositor>
  <registrant>American Scientific Publishing Group</registrant>
 </head>
 <body>
  <journal>
   <journal_metadata language="en">
    <full_title>Fusion: Practice and Applications</full_title>
    <abbrev_title>FPA</abbrev_title>
    <issn media_type="print">2770-0070</issn>
    <issn media_type="electronic">2692-4048</issn>
   </journal_metadata>
   <journal_issue>
    <publication_date media_type="online">
     <year>2021</year>
    </publication_date>
    <journal_volume>
     <volume>4</volume>
    </journal_volume>
    <issue>2</issue>
   </journal_issue>
   <journal_article publication_type="full_text">
    <titles>
     <title>Image Caption Generation and Comprehensive Comparison of Image Encoders</title>
    </titles>
    <contributors>
     <person_name sequence="first" contributor_role="author">
      <given_name>Shitiz</given_name>
      <surname>Gupta</surname>
      <affiliations>
       <institution>
        <institution_name>Bharati Vidyapeeth’s College of Engineering, New Delhi, India</institution_name>
       </institution>
      </affiliations>
     </person_name>
     <person_name sequence="additional" contributor_role="author">
      <given_name>Shubham</given_name>
      <surname>Agnihotri</surname>
      <affiliations>
       <institution>
        <institution_name>Bharati Vidyapeeth’s College of Engineering, New Delhi, India</institution_name>
       </institution>
      </affiliations>
     </person_name>
     <person_name sequence="additional" contributor_role="author">
      <given_name>Deepasha</given_name>
      <surname>Birla</surname>
      <affiliations>
       <institution>
        <institution_name>Bharati Vidyapeeth’s College of Engineering, New Delhi, India</institution_name>
       </institution>
      </affiliations>
     </person_name>
     <person_name sequence="additional" contributor_role="author">
      <given_name>Achin</given_name>
      <surname>Jain</surname>
      <affiliations>
       <institution>
        <institution_name>Bharati Vidyapeeth’s College of Engineering, New Delhi, India</institution_name>
       </institution>
      </affiliations>
     </person_name>
     <person_name sequence="additional" contributor_role="author">
      <given_name>Thavavel</given_name>
      <surname>Vaiyapuri</surname>
      <affiliations>
       <institution>
        <institution_name>College of computer engineering and sciences, Prince Sattam bin abdulaziz University, Saudi Arabia</institution_name>
       </institution>
      </affiliations>
     </person_name>
     <person_name sequence="additional" contributor_role="author">
      <given_name>Puneet Singh</given_name>
      <surname>Lamba</surname>
      <affiliations>
       <institution>
        <institution_name>Bharati Vidyapeeth’s College of Engineering, New Delhi, India</institution_name>
       </institution>
      </affiliations>
     </person_name>
    </contributors>
    <jats:abstract>
     <jats:p>Image caption generation is a stimulating multimodal task. Substantial advancements have been made in thefield of deep learning notably in computer vision and natural language processing. Yet, human-generated captions are still considered better, which makes it a challenging application for interactive machine learning. In this paper, we aim to compare different transfer learning techniques and develop a novel architecture to improve image captioning accuracy. We compute image feature vectors using different state-of-the-art transferlearning models which are fed into an Encoder-Decoder network based on Stacked LSTMs with soft attention,along with embedded text to generate high accuracy captions. We have compared these models on severalbenchmark datasets based on different evaluation metrics like BLEU and METEOR.</jats:p>
    </jats:abstract>
    <publication_date media_type="online">
     <year>2021</year>
    </publication_date>
    <pages>
     <first_page>42</first_page>
     <last_page>55</last_page>
    </pages>
    <publisher_item>
     <item_number item_number_type="article-number">730</item_number>
    </publisher_item>
    <ai:program name="AccessIndicators">
     <ai:license_ref applies_to="vor">https://creativecommons.org/licenses/by/4.0/</ai:license_ref>
    </ai:program>
    <doi_data>
     <doi>10.54216/FPA.040202</doi>
     <resource>https://www.americaspg.com/journal/3/article/730</resource>
     <collection property="text-mining">
      <item>
       <resource mime_type="application/pdf">https://www.americaspg.com/storage/7332.pdf</resource>
      </item>
     </collection>
    </doi_data>
   </journal_article>
  </journal>
 </body>
</doi_batch>
