<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "https://jats.nlm.nih.gov/publishing/1.3/JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xml:lang="ru">
  <front xmlns:xlink="http://www.w3.org/1999/xlink">
    <journal-meta>
      <journal-id journal-id-type="elibrary">9004</journal-id>
      <journal-title-group>
        <journal-title>Problems of information security. Computer systems</journal-title>
        <trans-title-group xml:lang="ru">
          <trans-title>Проблемы информационной безопасности. Компьютерные системы</trans-title>
        </trans-title-group>
      </journal-title-group>
      <issn pub-type="epub">2071-8217</issn>
    </journal-meta>
    <article-meta xmlns:xlink="http://www.w3.org/1999/xlink">
      <article-id pub-id-type="publisher-id">7</article-id>
      <title-group>
        <article-title>Research on methods for automated data extraction from web resources with dynamically changing content using AI agents</article-title>
        <trans-title-group xml:lang="ru">
          <trans-title>Исследование методов автоматизированного извлечения данных из веб-ресурсов с динамически изменяемым контентом с помощью ИИ-агентов</trans-title>
        </trans-title-group>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Abitov</surname>
            <given-names>Roman</given-names>
          </name>
          <xref ref-type="aff" rid="aff1"/>
          <email>abitov_roman@mail.ru</email>
        </contrib>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0000-0001-9862-1507</contrib-id>
          <name>
            <surname>Dakhnovich</surname>
            <given-names>Andrey</given-names>
          </name>
          <xref ref-type="aff" rid="aff1"/>
          <email>add@ibks.spbstu.ru</email>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Moskvin</surname>
            <given-names>Dmitry</given-names>
          </name>
          <xref ref-type="aff" rid="aff1"/>
          <email>moskvin_da@spbstu.ru</email>
        </contrib>
      </contrib-group>
      <aff id="aff1">Peter the Great St. Petersburg Polytechnic University</aff>
      <pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-10-09">
        <day>09</day>
        <month>10</month>
        <year>2026</year>
      </pub-date>
      <issue>3</issue>
      <fpage>96</fpage>
      <lpage>109</lpage>
      <self-uri xmlns:xlink="http://www.w3.org/1999/xlink" content-type="pdf" xlink:href="https://jisp.spbstu.ru/userfiles/images/oblozhki/3_2026.png"/>
      <abstract xml:lang="en">
        <p>This paper addresses automated extraction of data from web resources with dynamically changing content using AI agents. Data collection practice remains fragmented because of heterogeneous DOM structures, interactive user actions, anti-bot controls, and the predominance of unstructured web data, which increases maintenance costs of classical parsers. We propose a two-stage method that separates an expensive large language model analysis stage from an economical deterministic mass-collection stage. The method is implemented in a software prototype based on OpenClaw multi-agent orchestration, an isolated Chromium browser contour, the Model Context Protocol, and an authentication contour. We performed comparative analysis of scraping tools and large language models, classified feed types and active actions, and experimentally assessed extraction quality. On a reference dataset the primary quality metric reached F1 = 0.89. Limitations include selector fragility, evolving anti-bot mechanisms, and the computational cost of the configuration stage.</p>
      </abstract>
      <kwd-group xml:lang="en">
        <kwd>AI agents</kwd>
        <kwd>data extraction</kwd>
        <kwd>web scraping</kwd>
        <kwd>dynamic content</kwd>
        <kwd>large language models</kwd>
        <kwd>asynchronous processing</kwd>
        <kwd>fault tolerance</kwd>
        <kwd>data quality</kwd>
        <kwd>Model Context Protocol</kwd>
        <kwd>Chromium</kwd>
        <kwd>OpenClaw</kwd>
      </kwd-group>
    </article-meta>
  </front>
</article>
