2026-01-26 02:27:22 [scrapy.utils.log] INFO: Scrapy 2.14.1 started (bot: grabbers) 2026-01-26 02:27:22 [scrapy.utils.log] INFO: Versions: {'lxml': '6.0.2', 'libxml2': '2.14.6', 'cssselect': '1.3.0', 'parsel': '1.10.0', 'w3lib': '2.3.1', 'Twisted': '25.5.0', 'Python': '3.14.2 (main, Dec 6 2025, 14:00:34) [GCC 11.4.0]', 'pyOpenSSL': '25.3.0 (OpenSSL 3.5.4 30 Sep 2025)', 'cryptography': '46.0.3', 'Platform': 'Linux-6.12.63-84.121.amzn2023.x86_64-x86_64-with-glibc2.35'} 2026-01-26 02:27:22 [scrapy.addons] INFO: Enabled addons: [] 2026-01-26 02:27:22 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/extensions/feedexport.py:436: ScrapyDeprecationWarning: The `FEED_URI` and `FEED_FORMAT` settings have been deprecated in favor of the `FEEDS` setting. Please see the `FEEDS` setting docs for more details exporter = cls(crawler) 2026-01-26 02:27:22 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.logcount.LogCount', 'scrapy.extensions.memusage.MemoryUsage', 'scrapy.extensions.logstats.LogStats', 'grabbers.extensions.custom_fields.ExportCustomFieldsExtension', 'grabbers.extensions.kol_roles.ExportKolRolesExtension', 'grabbers.extensions.exporters.DefaultScrapedItemsExporter'] 2026-01-26 02:27:22 [scrapy.crawler] INFO: Overridden settings: {'BOT_NAME': 'grabbers', 'FEED_EXPORT_FIELDS': ['External Id', 'Parent External Id', 'Name', 'Plain text: Parent Name', 'Start date', 'Start time', 'End date', 'End time', 'Location', 'Plain text: Event Link', 'Plain text: Event Link (Excel)', 'Plain text: Description', 'Plain text: Speaker(s)', 'Plain text: Moderator(s)', 'Filter: Session Type'], 'FEED_FORMAT': 'xlsx', 'LOG_FILE': '/etc/scrapyd/dbs/logs/grabbers/ESGO ' '(events)/889075bcfa5e11f09c12dab4a64975df.log', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'grabbers.spiders', 'SPIDER_MODULES': ['grabbers.spiders'], 'TELNETCONSOLE_ENABLED': False, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64; rv:131.0) Gecko/20100101 ' 'Firefox/131.0'} 2026-01-26 02:27:22 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-26 02:27:22 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-26 02:27:22 [scrapy.middleware] INFO: Enabled downloader middlewares: ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware', 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware', 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware', 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware', 'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware', 'scrapy_fake_useragent.middleware.RetryUserAgentMiddleware', 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware', 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware', 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware', 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware', 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware', 'scrapy.downloadermiddlewares.stats.DownloaderStats'] 2026-01-26 02:27:22 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:44: ScrapyDeprecationWarning: RandomUserAgentMiddleware.process_request() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_request) 2026-01-26 02:27:22 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:47: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_response() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_response) 2026-01-26 02:27:22 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:50: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_exception() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_exception) 2026-01-26 02:27:22 [scrapy.middleware] INFO: Enabled spider middlewares: ['scrapy.spidermiddlewares.start.StartSpiderMiddleware', 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware', 'scrapy.spidermiddlewares.referer.RefererMiddleware', 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware', 'scrapy.spidermiddlewares.depth.DepthMiddleware'] 2026-01-26 02:27:23 [scrapy.middleware] INFO: Enabled item pipelines: ['grabbers.pipelines.CleanTextPipeline', 'grabbers.pipelines.DateTimeFormatPipeline', 'grabbers.pipelines.BoolFormatPipeline', 'grabbers.pipelines.ValidateEventNamePipeline', 'grabbers.pipelines.CountryCodePipeline', 'grabbers.pipelines.HyperlinkedEventLinkPipeline', 'grabbers.pipelines.ElasticSearchPipeline'] 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CleanTextPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: DateTimeFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: BoolFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ValidateEventNamePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CountryCodePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: HyperlinkedEventLinkPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:41: ScrapyDeprecationWarning: ElasticSearchPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.open_spider) 2026-01-26 02:27:23 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ElasticSearchPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-26 02:27:23 [scrapy.core.engine] INFO: Spider opened 2026-01-26 02:27:23 [ESGO (events)] ERROR: ElasticSearch client initialization failed. Make sure ElasticSearch is running and connection settings are correct. 2026-01-26 02:27:23 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2026-01-26 02:27:32 [ESGO (events)] WARNING: Cell Characters Limit Overflow. Event Name: Eposters External ID: 5239648 Field: speaker_s_plain 2026-01-26 02:27:33 [scrapy.core.engine] INFO: Closing spider (finished) 2026-01-26 02:27:33 [ESGO (events)] WARNING: Total fixed duplicated events: 2 2026-01-26 02:27:33 [ESGO (events)] WARNING: You can enable detailed logs for duplicated events, set `ENABLE_DUPLICATION_DETAILED_LOGS` environment variable to `1` 2026-01-26 02:27:33 [scrapy.utils.signal] ERROR: Error caught on signal handler: > Traceback (most recent call last): File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/utils/signal.py", line 192, in handler robustApply( ~~~~~~~~~~~^ receiver, signal=signal, sender=sender, *arguments, **named ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ), ^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/pydispatch/robustapply.py", line 55, in robustApply return receiver(*arguments, **named) File "/workspace/source/app/grabbers/extensions/exporters/mixins/delete_duplicated_subevents.py", line 284, in close_spider return super().close_spider(spider) ~~~~~~~~~~~~~~~~~~~~^^^^^^^^ File "/workspace/source/app/grabbers/extensions/exporters/mixins/add_missing_parent_names.py", line 65, in close_spider return super().close_spider(spider) ~~~~~~~~~~~~~~~~~~~~^^^^^^^^ File "/workspace/source/app/grabbers/extensions/exporters/mixins/fix_time_frames.py", line 250, in close_spider close_coroutine = super().close_spider(spider) File "/workspace/source/app/grabbers/extensions/exporters/base.py", line 32, in close_spider items = self.prepare_items(self.scrapped_items, spider) File "/workspace/source/app/grabbers/extensions/exporters/group.py", line 61, in prepare_items sub_events = self._prepare_sub_events(event, sub_events, spider) File "/workspace/source/app/grabbers/extensions/exporters/mixins/fix_time_frames.py", line 27, in _prepare_sub_events self._fix_time_frames(parent_event, sub_events, spider) ~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/workspace/source/app/grabbers/extensions/exporters/mixins/fix_time_frames.py", line 190, in _fix_time_frames event_link=parent_event_adapter["excel_link"], ~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/itemadapter/adapter.py", line 431, in __getitem__ return self.adapter.__getitem__(field_name) ~~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/itemadapter/adapter.py", line 243, in __getitem__ raise KeyError(field_name) KeyError: 'excel_link' 2026-01-26 02:27:33 [scrapy.statscollectors] INFO: Dumping Scrapy stats: {'downloader/request_bytes': 56464, 'downloader/request_count': 98, 'downloader/request_method_count/GET': 98, 'downloader/response_bytes': 2077108, 'downloader/response_count': 98, 'downloader/response_status_count/200': 98, 'elapsed_time_seconds': 9.100581, 'finish_reason': 'finished', 'finish_time': datetime.datetime(2026, 1, 26, 2, 27, 33, 56296, tzinfo=datetime.timezone.utc), 'item_scraped_count': 1933, 'items_per_minute': 12886.666666666668, 'log_count/INFO': 2, 'log_count/WARNING': 1, 'memusage/max': 229449728, 'memusage/startup': 229449728, 'request_depth_max': 1, 'response_received_count': 98, 'responses_per_minute': 653.3333333333334, 'scheduler/dequeued': 98, 'scheduler/dequeued/memory': 98, 'scheduler/enqueued': 98, 'scheduler/enqueued/memory': 98, 'start_time': datetime.datetime(2026, 1, 26, 2, 27, 23, 955715, tzinfo=datetime.timezone.utc)} 2026-01-26 02:27:33 [scrapy.core.engine] INFO: Spider closed (finished)