LumenData
Website:
lumendata.com
Job details:
Solution Architect – Databricks
Experience Required
- 10+ years of overall consulting / technology experience
- 7+ years of experience in Data Engineering, Data Platforms, Big Data, or Analytics
- Strong hands-on Databricks experience
- Minimum 6–8 end-to-end Databricks project implementations
Role Overview
We are looking for a highly experienced Senior FDE / Resident Solution Architect – Databricks who can work closely with enterprise customers to design, develop, optimize, and support scalable data engineering and analytics solutions on the Databricks platform.
The consultant should have strong hands-on expertise in Databricks, Apache Spark, PySpark, distributed computing, cloud platforms, performance optimization, and solution architecture.
This is a highly technical and client-facing role. The consultant should be able to independently drive architecture discussions, troubleshoot complex Databricks/Spark issues, provide implementation guidance, and support production deployments.
Key Responsibilities
- Design and implement scalable Databricks Lakehouse solutions.
- Work directly with customers to understand technical and business requirements.
- Define end-to-end data engineering and platform architecture.
- Build and optimize data pipelines using Databricks, Spark, PySpark, SQL, and Delta Lake.
- Design batch and streaming data-processing solutions.
- Provide technical guidance on Databricks architecture, development standards, and best practices.
- Troubleshoot complex Spark and Databricks performance issues.
- Optimize workloads for performance, scalability, reliability, and cost.
- Support enterprise Databricks platform implementation and modernization initiatives.
- Work with Databricks capabilities such as Delta Lake, Unity Catalog, Workflows, Auto Loader, Databricks SQL, Lakeflow/DLT, and Serverless.
- Design and support CI/CD processes for Databricks deployments.
- Work with DevOps and Infrastructure-as-Code tools such as Git, Terraform, Azure DevOps, GitHub, GitLab, or Jenkins.
- Provide guidance on Databricks security, governance, access control, and Unity Catalog.
- Support customer teams with architecture reviews, code reviews, troubleshooting, and technical mentoring.
- Work as a trusted technical advisor to customer architects, engineering teams, and stakeholders.
Mandatory Skills
Databricks
- Strong hands-on experience in Databricks development and architecture
- Minimum 6–8 Databricks projects delivered
- Delta Lake
- Databricks Lakehouse Architecture
- Unity Catalog
- Databricks Workflows
- Auto Loader
- Databricks SQL
- Batch and streaming workloads
- Cluster / compute configuration
- Performance optimization
Apache Spark / PySpark
- Strong hands-on Spark and PySpark development
- Deep understanding of Spark internals including:
- Driver and Executors
- DAG
- Jobs, Stages, and Tasks
- Partitioning
- Shuffle
- Memory management
- Catalyst Optimizer
- Adaptive Query Execution
- Spark SQL execution plans
- Data skew
- Join optimization
Data Engineering
- Strong ETL / ELT experience
- Data ingestion and transformation
- Data pipelines
- Data modeling
- Batch processing
- Streaming processing
- SQL
- Python / PySpark
- Large-scale distributed data processing
Cloud
Candidate must have:
- Deep expertise in at least one cloud platform: AWS / Azure / GCP
- Working knowledge of at least one additional cloud platform
Relevant cloud services may include:
AWS: S3, IAM, Glue, Lambda, Kinesis, Redshift
Azure: ADLS, ADF, Key Vault, Entra ID, Synapse, Event Hubs
GCP: GCS, BigQuery, Pub/Sub, Dataflow, IAM
Performance & Scalability
Strong experience with:
- Spark performance tuning
- Partitioning
- Shuffle optimization
- Data skew handling
- Join optimization
- Query optimization
- Delta table optimization
- Photon
- Cluster sizing
- Autoscaling
- Cost optimization
CI/CD & DevOps
Working knowledge of:
- Git
- CI/CD pipelines
- Terraform
- Databricks Asset Bundles
- Azure DevOps / GitHub / GitLab / Jenkins
- Dev / Test / UAT / Production deployment processes
MLOps
Working knowledge of:
- MLflow
- Experiment tracking
- Model Registry
- Model deployment / serving
- Model lifecycle management
Certification
Mandatory / Highly Preferred:
- Databricks Certified Data Engineer Professional
- Completion of relevant Databricks training/classes
Preferred Skills
- Enterprise Databricks architecture experience
- Databricks migrations / modernization
- Hadoop to Databricks migration
- Cloud data warehouse to Databricks migration
- Multi-cloud architecture exposure
- Unity Catalog implementation
- Terraform
- Databricks Asset Bundles
- MLflow / MLOps
- Data governance
- Streaming architecture
- Technical leadership / mentoring
Client-Facing Skills
The consultant must have strong experience in:
- Customer-facing technical discussions
- Architecture workshops
- Requirement gathering
- Solution design
- Technical presentations
- Stakeholder management
- Architecture reviews
- Technical recommendations
- Troubleshooting and problem solving
Ideal Candidate Profile
We are looking for candidates with approximately:
- 10–15+ years overall experience
- 7+ years Data Engineering / Big Data experience
- Strong Databricks and Spark experience
- 6–8+ Databricks project implementations
- Strong client-facing consulting background
- Databricks Data Engineer Professional certification
- Deep Spark / PySpark and Spark internals knowledge
- Strong performance tuning expertise
- Deep knowledge of one cloud platform and exposure to another
- Strong solution architecture and technical leadership capabilities
Vendor Submission Guidelines
Please submit only candidates who meet the core requirements. Each submission should clearly mention:
- Total Experience
- Relevant Data Engineering Experience
- Databricks Experience
- Number of Databricks Projects Delivered
- Spark / PySpark Experience
- Primary Cloud Expertise
- Secondary Cloud Exposure
- Databricks Certifications
- Current Location
- Current CTC / Rate
- Expected CTC / Rate
- Notice Period / Availability
- Current Organization
- Client-facing / Consulting Experience
- Brief summary of the candidate's strongest Databricks project
Click on Apply to know more.