2026-01-28 06:44:49 [scrapy.utils.log] INFO: Scrapy 2.14.1 started (bot: grabbers) 2026-01-28 06:44:49 [scrapy.utils.log] INFO: Versions: {'lxml': '6.0.2', 'libxml2': '2.14.6', 'cssselect': '1.3.0', 'parsel': '1.10.0', 'w3lib': '2.3.1', 'Twisted': '25.5.0', 'Python': '3.14.2 (main, Dec 6 2025, 14:00:34) [GCC 11.4.0]', 'pyOpenSSL': '25.3.0 (OpenSSL 3.5.4 30 Sep 2025)', 'cryptography': '46.0.3', 'Platform': 'Linux-6.12.63-84.121.amzn2023.x86_64-x86_64-with-glibc2.35'} 2026-01-28 06:44:49 [scrapy.addons] INFO: Enabled addons: [] 2026-01-28 06:44:49 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/extensions/feedexport.py:436: ScrapyDeprecationWarning: The `FEED_URI` and `FEED_FORMAT` settings have been deprecated in favor of the `FEEDS` setting. Please see the `FEEDS` setting docs for more details exporter = cls(crawler) 2026-01-28 06:44:50 [scrapy.middleware] INFO: Enabled extensions: ['scrapy.extensions.corestats.CoreStats', 'scrapy.extensions.logcount.LogCount', 'scrapy.extensions.memusage.MemoryUsage', 'scrapy.extensions.logstats.LogStats', 'scrapy.extensions.throttle.AutoThrottle', 'grabbers.extensions.kol_roles.ExportKolRolesExtension', 'grabbers.extensions.grabbers.abstract_online.AbstractOnlineExtension'] 2026-01-28 06:44:50 [scrapy.crawler] INFO: Overridden settings: {'AUTOTHROTTLE_ENABLED': True, 'AUTOTHROTTLE_MAX_DELAY': 5, 'AUTOTHROTTLE_TARGET_CONCURRENCY': 6, 'BOT_NAME': 'grabbers', 'FEED_EXPORT_FIELDS': ['External Id', 'Parent External Id', 'Name', 'Plain text: Parent Name', 'Start date', 'Start time', 'End date', 'End time', 'Location', 'Plain text: Event Link', 'Plain text: Event Link (Excel)', 'Filter: Session Type', 'Plain text: Session Number', 'Plain text: Description', 'Plain text: Authors', 'Filter: Keywords', 'Filter: Topic(s)', 'Plain text: Disclosures', 'Plain text: Abstract', 'Plain text: Presentation number', 'Plain text: Poster Number'], 'FEED_FORMAT': 'xlsx', 'LOG_FILE': '/etc/scrapyd/dbs/logs/grabbers/Actrims ' '(Posters)/d3814846fc1411f0aff536922b851908.log', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'grabbers.spiders', 'SPIDER_MODULES': ['grabbers.spiders'], 'TELNETCONSOLE_ENABLED': False, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64; rv:131.0) Gecko/20100101 ' 'Firefox/131.0'} 2026-01-28 06:44:50 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-28 06:44:50 [scrapy_fake_useragent.middleware] INFO: Using '' as the User-Agent provider 2026-01-28 06:44:50 [scrapy.middleware] INFO: Enabled downloader middlewares: ['scrapy.downloadermiddlewares.offsite.OffsiteMiddleware', 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware', 'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware', 'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware', 'scrapy_fake_useragent.middleware.RandomUserAgentMiddleware', 'scrapy_fake_useragent.middleware.RetryUserAgentMiddleware', 'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware', 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware', 'scrapy.downloadermiddlewares.redirect.RedirectMiddleware', 'scrapy.downloadermiddlewares.cookies.CookiesMiddleware', 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware', 'scrapy.downloadermiddlewares.stats.DownloaderStats'] 2026-01-28 06:44:50 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:44: ScrapyDeprecationWarning: RandomUserAgentMiddleware.process_request() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_request) 2026-01-28 06:44:50 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:47: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_response() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_response) 2026-01-28 06:44:50 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/core/downloader/middleware.py:50: ScrapyDeprecationWarning: RetryUserAgentMiddleware.process_exception() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(mw.process_exception) 2026-01-28 06:44:50 [scrapy.middleware] INFO: Enabled spider middlewares: ['scrapy.spidermiddlewares.start.StartSpiderMiddleware', 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware', 'scrapy.spidermiddlewares.referer.RefererMiddleware', 'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware', 'scrapy.spidermiddlewares.depth.DepthMiddleware'] 2026-01-28 06:44:51 [scrapy.middleware] INFO: Enabled item pipelines: ['grabbers.pipelines.CleanTextPipeline', 'grabbers.pipelines.DateTimeFormatPipeline', 'grabbers.pipelines.BoolFormatPipeline', 'grabbers.pipelines.ValidateEventNamePipeline', 'grabbers.pipelines.CountryCodePipeline', 'grabbers.pipelines.HyperlinkedEventLinkPipeline', 'grabbers.pipelines.ElasticSearchPipeline'] 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CleanTextPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: DateTimeFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: BoolFormatPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ValidateEventNamePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: CountryCodePipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: HyperlinkedEventLinkPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:41: ScrapyDeprecationWarning: ElasticSearchPipeline.open_spider() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.open_spider) 2026-01-28 06:44:51 [py.warnings] WARNING: /layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/pipelines/__init__.py:47: ScrapyDeprecationWarning: ElasticSearchPipeline.process_item() requires a spider argument, this is deprecated and the argument will not be passed in future Scrapy versions. If you need to access the spider instance you can save the crawler instance passed to from_crawler() and use its spider attribute. self._check_mw_method_spider_arg(pipe.process_item) 2026-01-28 06:44:51 [scrapy.core.engine] INFO: Spider opened 2026-01-28 06:44:51 [Actrims (Posters)] ERROR: ElasticSearch client initialization failed. Make sure ElasticSearch is running and connection settings are correct. 2026-01-28 06:44:51 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2026-01-28 06:44:57 [Actrims (Posters)] INFO: Found days: Feb05 Feb06 Feb07 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Patient Demographics and Epidemiology'} id: 927 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Non-MS Neuroimmune Conditions'} id: 327 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Imaging'} id: 763 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Disease Mechanisms and Pathogenesis'} id: 878 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Diagnosis and Differential Diagnosis'} id: 925 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Clinical Measures and Assessment'} id: 924 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Digital Tools'} id: 926 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Session Type', 'filter_choice': 'Partner Organizations'} id: 6 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Clinical Trials'} id: 750 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Session Type', 'filter_choice': 'Cutting Edge Developments'} id: 4 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Topic', 'filter_choice': 'Biomarkers and Metabolomics'} id: 33 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Sessions for day -> day_id: 3 | attempt: 0 | event count: 4 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Search for {'filter_name': 'Session Type', 'filter_choice': 'Poster Session'} id: 7 2026-01-28 06:45:09 [Actrims (Posters)] INFO: Sessions for day -> day_id: 2 | attempt: 0 | event count: 5 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search for {'filter_name': 'Session Type', 'filter_choice': 'Invited Session'} id: 5 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 927[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Sessions for day -> day_id: 1 | attempt: 0 | event count: 6 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 327[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 763[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 878[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 925[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 924[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 926[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 6[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 750[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 4[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 33[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 7[0]: Complete 2026-01-28 06:45:10 [Actrims (Posters)] INFO: Search 5[0]: Complete 2026-01-28 06:45:20 [scrapy.core.engine] INFO: Closing spider (finished) 2026-01-28 06:45:21 [Actrims (Posters)] INFO: Save the actual congress fields. 2026-01-28 06:45:21 [scrapy.utils.signal] ERROR: Error caught on signal handler: > Traceback (most recent call last): File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/scrapy/utils/signal.py", line 192, in handler robustApply( ~~~~~~~~~~~^ receiver, signal=signal, sender=sender, *arguments, **named ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ), ^ File "/layers/paketo-buildpacks_pip-install/packages/lib/python3.14/site-packages/pydispatch/robustapply.py", line 55, in robustApply return receiver(*arguments, **named) File "/workspace/source/app/grabbers/pipelines/store_to_elasticsearch.py", line 162, in spider_closed self.client.add_grabber_fields( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ AttributeError: 'NoneType' object has no attribute 'add_grabber_fields' 2026-01-28 06:45:21 [scrapy.extensions.feedexport] INFO: Stored xlsx feed (418 items) in: s3://usummit-prod-grabbers-bucket/Actrims (Posters)/Actrims (Posters)-2101_2026-01-28_06-44-42.xlsx 2026-01-28 06:45:21 [scrapy.statscollectors] INFO: Dumping Scrapy stats: {'downloader/request_bytes': 69846, 'downloader/request_count': 90, 'downloader/request_method_count/GET': 73, 'downloader/request_method_count/POST': 17, 'downloader/response_bytes': 2432193, 'downloader/response_count': 90, 'downloader/response_status_count/200': 89, 'downloader/response_status_count/201': 1, 'elapsed_time_seconds': 29.064678, 'feedexport/success_count/S3FeedStorage': 1, 'finish_reason': 'finished', 'finish_time': datetime.datetime(2026, 1, 28, 6, 45, 20, 864471, tzinfo=datetime.timezone.utc), 'item_scraped_count': 477, 'items_per_minute': 986.8965517241379, 'log_count/INFO': 32, 'memusage/max': 229658624, 'memusage/startup': 229658624, 'request_depth_max': 5, 'response_received_count': 90, 'responses_per_minute': 186.20689655172413, 'scheduler/dequeued': 90, 'scheduler/dequeued/memory': 90, 'scheduler/enqueued': 90, 'scheduler/enqueued/memory': 90, 'start_time': datetime.datetime(2026, 1, 28, 6, 44, 51, 799793, tzinfo=datetime.timezone.utc)} 2026-01-28 06:45:21 [scrapy.core.engine] INFO: Spider closed (finished)