Website:
ms.holdings
Job details:
About the job
Associate - Python Developer
(Data Collection & Pipelines) 2+ years' experience.
About the Role:
We are building a data collection and storage infrastructure from scratch - and we are looking for a Python Developer to
own it. You will integrate with third-party data APIs, build custom web crawlers where public sources do not offer an API,
design the database structure, and normalise data from multiple sources into a clean, consistent store ready for analysis.
Once the infrastructure is operational, the scope grows. As the platform matures, there is a clear path toward building
lightweight internal tooling on top of the data layer you have built. This is a long-term role with growing ownership.
What You Will Do:
— Integrate with multiple third-party data APIs - handle authentication, rate limiting, pagination, retries, and
inconsistent response schemas
— Build custom web crawlers where public data sources do not offer an API
— Design and manage a relational database to store, organise, and retrieve collected data reliably
— Build standardisation and normalisation logic that transforms raw data from different sources and formats into a
clean, consistent internal structure
— Handle deduplication and data quality management across all sources
— Build scheduled and on-demand data collection pipelines - automated runs on a defined cadence plus manual trigger
capability
— Write clean, well-documented code
— Maintain the data infrastructure on an ongoing basis
Required Skills:
— Python - 2+ years as primary language. Strong, not just familiar.
— REST API integration - real experience consuming multiple third-party APIs and handling practical integration
challenges
— Web crawling / scraping - has built custom crawlers using Scrapy, BeautifulSoup, Playwright, or similar. Understands
dynamic content, anti-bot handling, and responsible crawling
— Relational Database - schema design from scratch, query writing, indexing, and ongoing management
— Data normalisation and deduplication across multiple sources and formats
— Scheduling and automation - APScheduler, cron, or equivalent
— Cloud Basics - comfortable hosting scheduled jobs and managing a cloud database (AWS, GCP, or Azure)
Preferred Skills:
— Experience at a data product company, data infrastructure firm, media monitoring firm, data aggregator, or fintech
where data collection was core work.
— FastAPI or Flask
— Exposure to multilingual or non-English text data and OSINT frameworks
— Basic frontend experience - as the platform grows, there will be opportunity to build lightweight internal interfaces on
top of the data layer
Who We Are Looking For:
Someone who takes full ownership of what they build. You have built a data collection pipeline from scratch before - APIs,
crawlers, or both. You design clean schemas, write readable code, document your decisions. You are comfortable working
independently and bringing sound technical judgment to the problems in front of you.
If you've previously built API integrations, crawlers, or internal data platforms and you want to build infrastructure that
becomes the backbone of an intelligence product - we would like to hear from you.
Work Location- Kinfra Technopark Kakkanchery (onsite)
Working hours - 9.30AM to 6.30AM
Days- Mondays - Friday (fixed off on saturday and sunday)
you can drop your resumes on aswanth.dinesh@ms-ca.com
Click on Apply to know more.