Role Overview
Design, develop, and maintain scalable AI-powered data ingestion and processing platforms using Python. Build intelligent data pipelines, web crawling solutions, and AI/ML services that process structured and unstructured data while ensuring scalability, performance, and reliability.
About CUBE
CUBE is a global RegTech business defining and implementing the gold standard of regulatory intelligence for the financial services industry. We deliver our services through intuitive SaaS solutions, powered by AI, to simplify the complex and everchanging world of compliance for our clients.
Key Responsibilities
- Design and develop web crawling, data ingestion, and AI-powered processing solutions using Python.
- Build scalable data pipelines for processing HTML, PDF, and other structured/unstructured data sources.
- Develop REST APIs to support AI/ML services and data processing workflows.
- Implement AI/ML models or integrate LLMs for document parsing, classification, summarization, and information extraction.
- Design and optimize background processing, task scheduling, and distributed workloads.
Requirements & Eligibility
- 2-3 years of Python development experience with FastAPI, Django, or similar frameworks.
- Strong experience in AI/ML, Generative AI, or LLM application development.
- Experience with web scraping frameworks such as Scrapy or Playwright.
- Hands-on experience with NLP, document processing, or information extraction techniques.
- Experience integrating LLMs (OpenAI, Claude, Gemini, Llama, or similar) using frameworks like LangChain or LlamaIndex.
- Strong knowledge of REST APIs, asynchronous programming, and distributed task processing (Celery, RabbitMQ, Kafka).
Required Skills & Tech
PythonFastAPIDjangoScrapyPlaywright