2026-01-26 05:59:20 [scrapy.utils.log] INFO: Scrapy 2.14.1 started (bot: grabbers) 2026-01-26 05:59:20 [scrapy.utils.log] INFO: Versions: {'lxml': '6.0.2', 'libxml2': '2.14.6', 'cssselect': '1.3.0', 'parsel': '1.10.0', 'w3lib': '2.3.1', 'Twisted': '25.5.0', 'Python': '3.14.2 (main, Dec 6 2025, 14:00:34) [GCC 11.4.0]', 'pyOpenSSL': '25.3.0 (OpenSSL 3.5.4 30 Sep 2025)', 'cryptography': '46.0.3', 'Platform': 'Linux-6.12.58-82.121.amzn2023.x86_64-x86_64-with-glibc2.35'} 2026-01-26 05:59:20 [scrapy.addons] INFO: Enabled addons: [] 2026-01-26 05:59:20 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/extensions/feedexport.py:436: ScrapyDeprecationWarning: The `FEED_URI` and `FEED_FORMAT` settings have been deprecated in favor of the `FEEDS` setting. Please see the `FEEDS` setting docs for more details exporter = cls(crawler) 2026-01-26 05:59:20 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.logcount.LogCount', 'scrapy.extensions.memusage.MemoryUsage', 'scrapy.extensions.logstats.LogStats', 'grabbers.extensions.custom_fields.ExportCustomFieldsExtension', 'grabbers.extensions.kol_roles.ExportKolRolesExtension', 'grabbers.extensions.exporters.DefaultScrapedItemsExporter'] 2026-01-26 05:59:20 [scrapy.crawler] INFO: Overridden settings: {'BOT_NAME': 'grabbers', 'FEED_EXPORT_FIELDS': ['External Id', 'Parent External Id', 'Name', 'Plain text: Parent Name', 'Start date', 'Start time', 'End date', 'End time', 'Location', 'Plain text: Event Link', 'Plain text: Event Link (Excel)', 'Plain text: Description', 'Plain text: Speaker(s)', 'Plain text: Moderator(s)', 'Filter: Session Type'], 'FEED_FORMAT': 'xlsx', 'LOG_FILE': '/etc/scrapyd/dbs/logs/grabbers/ESGO ' '(events)/23325cc6fa7c11f083530a943f6e3bcc.log', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'grabbers.spiders', 'SPIDER_MODULES': ['grabbers.spiders'], 'TELNETCONSOLE_ENABLED': False, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64; rv:131.0) Gecko/20100101 ' 'Firefox/131.0'} 2026-01-26 05:59:20 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-26 05:59:20 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-26 05:59:20 [scrapy.middleware] INFO: Enabled downloader middlewares: ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware', 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware', 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware', 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware', 'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware', 'scrapy_fake_useragent.middleware.RetryUserAgentMiddleware', 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware', 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware', 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware', 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware', 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware', 'scrapy.downloadermiddlewares.stats.DownloaderStats'] 2026-01-26 05:59:20 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:44: ScrapyDeprecationWarning: RandomUserAgentMiddleware.process_request() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_request) 2026-01-26 05:59:20 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:47: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_response() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_response) 2026-01-26 05:59:20 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:50: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_exception() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_exception) 2026-01-26 05:59:20 [scrapy.middleware] INFO: Enabled spider middlewares: ['scrapy.spidermiddlewares.start.StartSpiderMiddleware', 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware', 'scrapy.spidermiddlewares.referer.RefererMiddleware', 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware', 'scrapy.spidermiddlewares.depth.DepthMiddleware'] 2026-01-26 05:59:21 [scrapy.middleware] INFO: Enabled item pipelines: ['grabbers.pipelines.CleanTextPipeline', 'grabbers.pipelines.DateTimeFormatPipeline', 'grabbers.pipelines.BoolFormatPipeline', 'grabbers.pipelines.ValidateEventNamePipeline', 'grabbers.pipelines.CountryCodePipeline', 'grabbers.pipelines.HyperlinkedEventLinkPipeline', 'grabbers.pipelines.ElasticSearchPipeline'] 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CleanTextPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: DateTimeFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: BoolFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ValidateEventNamePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CountryCodePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: HyperlinkedEventLinkPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:41: ScrapyDeprecationWarning: ElasticSearchPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.open_spider) 2026-01-26 05:59:21 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ElasticSearchPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 05:59:21 [scrapy.core.engine] INFO: Spider opened 2026-01-26 05:59:21 [ESGO (events)] ERROR: ElasticSearch client initialization failed. Make sure ElasticSearch is running and connection settings are correct. 2026-01-26 05:59:21 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2026-01-26 05:59:38 [ESGO (events)] WARNING: Cell Characters Limit Overflow. Event Name: Eposters External ID: 5239648 Field: speaker_s_plain 2026-01-26 05:59:38 [scrapy.core.engine] INFO: Closing spider (finished) 2026-01-26 05:59:38 [ESGO (events)] WARNING: Total fixed duplicated events: 2 2026-01-26 05:59:38 [ESGO (events)] WARNING: You can enable detailed logs for duplicated events, set `ENABLE_DUPLICATION_DETAILED_LOGS` environment variable to `1` 2026-01-26 05:59:39 [ESGO (events)] WARNING: Event time frame was changed to fit sub events: External id: 5239656 end time was changed from 17:15 to 17:35 2026-01-26 05:59:39 [ESGO (events)] WARNING: Please, do not forget to add links to sub events with fixed times to task comments. 2026-01-26 05:59:40 [scrapy.extensions.feedexport] INFO: Stored xlsx feed (1931 items) in: s3://usummit-prod-grabbers-bucket/ESGO (events)/ESGO (events)-2092_2026-01-26_05-59-12.xlsx 2026-01-26 05:59:40 [scrapy.statscollectors] INFO: Dumping Scrapy stats: {'downloader/request_bytes': 56043, 'downloader/request_count': 98, 'downloader/request_method_count/GET': 98, 'downloader/response_bytes': 2077108, 'downloader/response_count': 98, 'downloader/response_status_count/200': 98, 'elapsed_time_seconds': 17.127077, 'feedexport/success_count/S3FeedStorage': 1, 'finish_reason': 'finished', 'finish_time': datetime.datetime(2026, 1, 26, 5, 59, 38, 873603, tzinfo=datetime.timezone.utc), 'item_scraped_count': 1933, 'items_per_minute': 6822.352941176471, 'log_count/INFO': 2, 'log_count/WARNING': 1, 'memusage/max': 229117952, 'memusage/startup': 229117952, 'request_depth_max': 1, 'response_received_count': 98, 'responses_per_minute': 345.88235294117646, 'scheduler/dequeued': 98, 'scheduler/dequeued/memory': 98, 'scheduler/enqueued': 98, 'scheduler/enqueued/memory': 98, 'start_time': datetime.datetime(2026, 1, 26, 5, 59, 21, 746526, tzinfo=datetime.timezone.utc)} 2026-01-26 05:59:40 [scrapy.core.engine] INFO: Spider closed (finished)