Data Architecture
Data models, source-of-truth design, lifecycle state, reconciliation, historical data, identifiers, and system boundaries.
Data Architecture · Data Engineering · Software Systems
I design and build production data platforms, automated processing pipelines, cloud infrastructure, search and e-commerce data systems, and bespoke software where correctness, scale, and operational reliability matter.
Capabilities
Systems are designed around authoritative data, predictable state, verification, and maintainability.
Data models, source-of-truth design, lifecycle state, reconciliation, historical data, identifiers, and system boundaries.
High-volume ingestion, ETL and ELT, transformation, validation, parallel processing, and repeatable batch pipelines.
Production Java and Spring Boot applications, APIs, database services, automation tooling, and long-running processing systems.
AWS infrastructure, S3 and CDN publication, containerized workloads, CI/CD, deployment automation, verification, and cleanup.
Product information, taxonomy, facets, variants, feeds, structured product data, search infrastructure, and technical SEO.
XML, XSLT, XProc, DITA-OT, RDF, SPARQL, Schema.org, JSON-LD, and publishing pipelines.
Selected Work
A selection of recent and specialist engineering work.
Backend architecture and engineering supporting a large B2B e-commerce website with an excellent store rating, processing hundreds of thousands of products across supplier data, product information, taxonomy, variants, search feeds, structured data, images, and web publication.
Automated ingestion, image identity tracking, WebP conversion, resizing, persistent processing state, alternate-image lifecycle handling, CDN publication, and post-upload verification with failure recovery.
Parallel acquisition and transformation of historical and daily SEC data, including structured extraction from XML, HTML, and unstructured filings for targeted downstream analysis.
CI/CD and document-engineering systems that transform structured content into web and print outputs using XML technologies, semantic data, containerized tooling, and automated publishing workflows.
Technology
Java, Python, Bash, SQL, XSLT, XML/XQuery, HTML, CSS
Spring Boot, Spring Data, MongoDB, MySQL, REST APIs, Hibernate
AWS, S3, CloudFront, EC2, CloudFormation, Docker, Linux, GitLab CI/CD
ETL/ELT, parallel processing, validation, reconciliation, web data ingestion, data pipelines
Schema.org, JSON-LD, taxonomy, faceted search, product feeds, semantic HTML, technical SEO
DITA-OT, XSLT, XProc, RDF, SPARQL, XML Schema, Schematron, MarkLogic, GraphDB, Jena
Experience
I have developed software and data solutions since 2002 for organizations and projects across North America, Europe, Africa, and Australia. My work has included publishing, financial and regulatory data, web platforms, infrastructure, analytics, semantic systems, and e-commerce.
Earlier specialist work includes rough-set data analysis using Binary Decision Diagrams and a publication in Revista Real Academia de Ciencias.
Oxford University Press, Dorling Kindersley, Pearson Education, Penguin Books, European Union OHIM, The Weather Network, Brock University, financial and regulatory-data projects, and international consulting engagements.
Contact
Available for architecture, engineering, technical review, and bespoke system development.