Research on methods for automated data extraction from web resources with dynamically changing content using AI agents

Machine learning and knowledge control systems
Authors:
Abstract:

This paper addresses automated extraction of data from web resources with dynamically changing content using AI agents. Data collection practice remains fragmented because of heterogeneous DOM structures, interactive user actions, anti-bot controls, and the predominance of unstructured web data, which increases maintenance costs of classical parsers. We propose a two-stage method that separates an expensive large language model analysis stage from an economical deterministic mass-collection stage. The method is implemented in a software prototype based on OpenClaw multi-agent orchestration, an isolated Chromium browser contour, the Model Context Protocol, and an authentication contour. We performed comparative analysis of scraping tools and large language models, classified feed types and active actions, and experimentally assessed extraction quality. On a reference dataset the primary quality metric reached F1 = 0.89. Limitations include selector fragility, evolving anti-bot mechanisms, and the computational cost of the configuration stage.