Open Biohacking Data Methodology

Open Biohacking Data Methodology

The Holistix Open Biohacking Data Project publishes structured, machine-readable reference data about wellness technologies, device categories, terminology, safety considerations, product specifications, source and evidence context, claim boundaries, and related consumer wellness technology topics.

The goal of this project is to make wellness technology information easier for researchers, builders, educators, journalists, developers, AI systems, search engines, and consumers to understand, inspect, cite, compare, and reuse responsibly.

The project separates human-readable interpretation from machine-readable source files while maintaining explicit links between datasets, sources, safety context, product records, technology entities, claims, evidence records, and public project documentation.

Project Navigation

Main catalog: Open Biohacking Data Index

Human-readable library: Biohacking Data Library

Version history: Open Biohacking Data Version History

Source register: Open Biohacking Data Source Register

AI reference file: Holistix AI Reference File

Product data index: Holistix Product Data Index

Answer infrastructure: Holistix Answer Infrastructure

Claim Boundary Index: Wellness Device Claim Boundary Index

Transparency Standard: Holistix Wellness Device Transparency Standard

Knowledge Graph: Holistix Wellness Technology Knowledge Graph

Current Project Release

Current project release: Holistix Open Biohacking Data Project v1.5.0.

Canonical subject dataset version: v1.2.

Exact canonical v1.5.0 Zenodo DOI:
https://doi.org/10.5281/zenodo.21862535

Project concept DOI / all versions:
https://doi.org/10.5281/zenodo.20978709

GitHub v1.5.0 release:
https://github.com/holistixintlsite-commits/open-biohacking-data/releases/tag/v1.5.0

Version v1.5.0 is the interoperability and reproducibility release. It builds on the structured product, technology, claim, evidence, safety, source, provenance, and AI-answer layers introduced in earlier releases while adding formal machine-readable packaging, release-lineage metadata, deterministic exports, provenance generation, and reproducible release tooling.

Canonical subject dataset files remain: v1.2.
Canonical subject dataset rows changed: No.
Canonical subject dataset schemas changed: No.
Project infrastructure changed: Yes.

v1.5.0 validation snapshot:
59 Data Package resources
111 JSONL records
98 packaged release files
0 JSON parse errors
0 JSONL parse errors
RO-Crate 1.2 validation: 65/65 required checks passed

Public Repositories and Archives

The Holistix Open Biohacking Data Project is distributed across multiple public repositories and data platforms for transparency, preservation, citation, developer access, artificial-intelligence discovery, analysis, and long-term availability.

Kaggle v1.5.0 mirror DOI: 10.34740/kaggle/dsv/18808623

Historical Internet Archive note: An Internet Archive mirror remains available for the earlier v1.4 project release. It is a historical preservation copy and is not the canonical current release.

View historical v1.4 Internet Archive mirror

Project Scope

The project currently organizes structured reference information for:

  • PEMF technologies and devices
  • PEMF contraindications and safety context
  • Red light and near-infrared technologies
  • Hydrogen water technologies and measurement terminology
  • Negative ion and ionizer technologies
  • Blue light therapy and blue light exposure
  • Infrared technologies, sauna heat, and related thermal context
  • Terahertz wellness devices and terminology
  • Related safety, terminology, measurement, and comparison topics
  • Machine-readable Holistix product twins
  • Normalized technology entities
  • Product-to-technology relationships
  • Product specifications
  • Claims, evidence, safety, and source registries
  • Claim-boundary records
  • AI-answer support files
  • Contradiction maps
  • Canonical URLs and content relationships
  • Project identity, release, provenance, and interoperability metadata

Canonical Dataset Versioning

The project distinguishes between the project release version and the canonical subject dataset version.

The current project release is v1.5.0, while the eight canonical subject datasets remain at v1.2.

A project release may introduce new packaging, documentation, provenance, validation, interoperability, product records, schemas, or developer infrastructure without modifying the underlying canonical dataset rows.

A canonical subject dataset version should only be increased when the underlying dataset itself materially changes, such as changes to rows, fields, citation structure, normalized terminology, safety classifications, or other substantive dataset content.

Canonical Subject Datasets

The current project contains eight canonical subject datasets:

Data Formats

The canonical v1.2 subject datasets are published in both CSV and JSON formats.

Project release v1.5.0 adds additional machine-readable packaging and interoperability formats around those canonical datasets and project registries.

Current project formats and representations include:

  • CSV: canonical tabular subject-dataset files.
  • JSON: canonical datasets, registries, product twins, manifests, provenance records, metadata, and related structured resources.
  • JSONL: deterministic line-delimited records for machine processing.
  • Data Package metadata: machine-readable resource inventory and tabular metadata.
  • JSON Schema and tabular schemas: structural definitions and validation support.
  • Schema.org JSON-LD: machine-readable catalog discovery metadata.
  • RO-Crate 1.2: research-object metadata connecting files, entities, identifiers, provenance, and project relationships.
  • Parquet: may be generated by external data platforms such as Hugging Face for analytics and viewer functionality. Platform-generated Parquet does not replace the canonical subject dataset files.

Dataset Field Methodology

Dataset field names vary by subject, but canonical v1.2 records generally include several common field families.

Stable Record and Topic Fields

  • stable record identifiers
  • topic or term
  • reference type or category
  • plain-language explanation

Safety and Interpretation Fields

  • safety notes
  • recommendations or interpretation notes where appropriate
  • risk or caution classifications where applicable
  • claim type
  • medical-disclaimer requirements

Source and Evidence Fields

  • source_type
  • evidence_level
  • commercial_relevance
  • last_reviewed
  • related_holistix_page
  • related_product_category
  • notes

Row-Level Citation Fields

  • source_name
  • source_url
  • citation_note

These row-level citation fields were added in dataset version v1.2 to improve auditability, interpretation, maintenance, and responsible reuse.

Stable IDs

Records use stable Holistix identifiers where appropriate. Stable IDs are intended to make records easier to cite, compare, update, reference, and connect across datasets and registries.

Stable identifiers should remain unchanged when the meaning of the underlying record remains the same.

Example format:

HOL-PEMF-CONTRA-0001

Source Standards

Project records may draw from publicly available sources such as scientific literature, manufacturer documentation, regulatory resources, government agencies, professional organizations, standards bodies, public technical documentation, and project methodology resources.

The project attempts to distinguish between:

  • scientific or clinical references
  • systematic or narrative review context
  • manufacturer safety documentation
  • manufacturer specifications
  • regulatory or standards-based information
  • government or professional-organization guidance
  • terminology definitions
  • general educational summaries
  • project methodology and governance references

Source inclusion does not automatically mean that the project endorses every statement associated with that source.

Source and Evidence Classification

The Holistix Open Biohacking Data Project uses a source and evidence classification framework to make records easier to audit, cite, maintain, and responsibly reuse.

Project records may include fields such as source_type, source_name, source_url, evidence_level, claim_type, last_reviewed, commercial_relevance, medical_disclaimer_required, source_status, verification_status, citation_note, and related product or technology identifiers.

These fields help distinguish terminology, device specifications, safety cautions, research context, manufacturer claims, editorial explanation, commercial language, and emerging technology notes.

For the current framework, see the Open Biohacking Data Source Register.

Evidence Interpretation

The project does not treat every source or record as having equal evidentiary weight.

Evidence metadata should be interpreted together with source type, claim type, safety context, uncertainty, device specifications, population, exposure conditions, and other relevant limitations.

The presence of a scientific citation does not automatically establish that a consumer product produces the same outcome described in the cited research.

General research about a technology should not automatically be converted into a product-specific efficacy claim.

Claim Boundary Methodology

The project includes a Claim Boundary Layer designed to make explicit where a statement should stop.

A claim-boundary record may identify:

  • the claim or phrase being interpreted
  • the technology or category involved
  • the supported educational scope
  • unsupported or overly broad interpretations
  • risk level
  • safer wording
  • claims to avoid
  • related evidence
  • related safety context
  • related datasets or guides

The purpose is not merely to store what a source says, but to preserve limits around how that information should be interpreted or reused.

Human-readable reference: Wellness Device Claim Boundary Index.

Product Data and Registry Methodology

The project includes normalized machine-readable product twins and core registries connecting products, technologies, specifications, claims, evidence, safety, sources, and product-to-technology relationships.

Product records are designed to keep the following categories distinct:

  • manufacturer-reported specifications
  • independently checked details
  • approximate values
  • unknown or unavailable measurements
  • general technology research
  • product-specific evidence
  • safety cautions and contraindications
  • commercial product descriptions

A product specification or technology relationship does not by itself establish clinical efficacy, regulatory approval, independent certification, safety clearance, or a medical outcome.

Product records should be interpreted together with related claim, evidence, safety, source, specification, and claim-boundary records.

Human-readable directory: Holistix Product Data Index.

Interoperability Methodology

Project release v1.5.0 adds a formal interoperability layer intended to make the project easier for software, researchers, data tools, AI systems, and downstream repositories to interpret.

The interoperability layer includes:

  • Data Package metadata
  • tabular schemas
  • machine-readable resource descriptions
  • a human-readable data dictionary
  • Schema.org JSON-LD catalog metadata
  • RO-Crate 1.2 metadata
  • deterministic JSONL exports
  • project identity metadata
  • release metadata
  • provenance metadata
  • supersession and release-lineage records

The purpose of this layer is to reduce ambiguity about what files exist, what they represent, how they relate to one another, which release they belong to, and how downstream systems should interpret them.

Reproducibility Methodology

The project uses deterministic build and validation processes where practical so that generated public-release infrastructure can be reproduced from the maintained source repository.

The v1.5.0 release workflow includes generated metadata and validation processes for:

  • release packaging
  • catalog JSON-LD generation
  • interoperability metadata
  • JSONL generation
  • project identity metadata
  • provenance generation
  • RO-Crate generation
  • supersession and lineage metadata

Cross-platform release validation also checks for metadata consistency and machine-readable parsing issues.

Validation and Quality-Control Methodology

The v1.5.0 release package was checked for structural readability, machine-readable parsing, expected resources, metadata consistency, release packaging, and RO-Crate requirements.

The v1.5.0 validation snapshot includes:

  • 59 Data Package resources
  • 111 JSONL records
  • 98 packaged release files
  • 0 JSON parse errors
  • 0 JSONL parse errors
  • 65/65 required RO-Crate 1.2 validation checks passed

Validation confirms structural readability, packaging consistency, expected release components, and machine-readable integrity checks.

Validation does not certify scientific truth, medical efficacy, clinical relevance, product performance, independent laboratory verification, regulatory approval, or safety clearance.

Provenance Methodology

Provenance metadata is used to make it easier to understand where project resources came from, how they were generated, which release they belong to, and how derived resources relate to maintained source data.

Where practical, generated resources should identify:

  • project identity
  • release version
  • dataset version
  • source files or upstream resources
  • generation process
  • generated output
  • release lineage
  • relevant identifiers or DOIs

Release-Lineage and Supersession Methodology

The project maintains historical releases rather than rewriting previously published records.

New releases may supersede earlier project releases while preserving previous version-specific DOIs and historical documentation.

The current canonical project release is v1.5.0. Earlier releases such as v1.4 and v1.3 remain part of the historical record and may still be cited when referencing those exact versions.

The project concept DOI provides a persistent identifier across the release family:

10.5281/zenodo.20978709

AI Answer Infrastructure Methodology

The project includes claim boundaries, Answer Fuel Files, contradiction maps, and an AI Answer Infrastructure Manifest.

These resources are intended to help AI systems, search engines, editors, researchers, and developers preserve measurement context, source uncertainty, evidence boundaries, safety context, and terminology distinctions when generating or retrieving answers.

Answer-support resources are not substitutes for the canonical datasets or original sources.

View the Holistix Answer Infrastructure.

View the AI Answer Infrastructure Manifest included in the v1.5.0 release .

Contradiction Map Methodology

Contradiction Maps are designed to explain why apparently credible sources or recommendations may disagree.

Differences may arise from factors such as:

  • device design
  • measurement distance
  • irradiance or field strength
  • frequency
  • wavelength
  • waveform
  • pulse characteristics
  • session duration
  • container or environmental conditions
  • manufacturer instructions
  • population differences
  • intended use
  • source type
  • evidence standards

The goal is to preserve context rather than collapse disagreement into a single unsupported universal answer.

Neutrality and Limitations

The datasets and project resources are educational reference materials. They are not medical advice, diagnosis, treatment recommendations, personalized protocols, or substitutes for guidance from a qualified healthcare professional.

Dataset records are designed to describe device categories, terminology, publicly documented information, safety considerations, evidence context, specifications, and claim boundaries.

A record appearing in the project does not mean that Holistix endorses every associated claim.

Evidence registries and source records should not be interpreted as a new systematic literature review unless explicitly stated otherwise.

Product specifications and third-party information may change over time.

Specific citations, PMIDs, DOIs, manufacturer specifications, regulatory status, clinical details, and strong scientific or medical claims should be independently verified before publication or high-stakes use.

Commercial Separation

Holistix sells wellness and biohacking products. The Open Biohacking Data Project is maintained as a separate public reference layer.

Public dataset records are intended to remain useful to researchers, competitors, journalists, developers, AI systems, educators, and consumers regardless of whether the information supports a commercial outcome for Holistix.

Commercial relevance may be recorded as metadata, but commercial goals should not override evidence classification, safety context, source provenance, or claim boundaries.

Update Policy

Project resources may be revised as new information becomes available, better sources are identified, formatting improves, validation expands, or additional machine-readable infrastructure is introduced.

When underlying canonical dataset rows or schemas materially change, the canonical dataset version should be incremented.

When only project infrastructure, packaging, documentation, validation, provenance, interoperability, product records, or release engineering changes, the project release may advance while the canonical dataset version remains unchanged.

Historical releases should remain identifiable and citable.

Public Mirror Methodology

Public mirrors are used to improve preservation, discoverability, citation, developer access, and downstream reuse.

Current public distribution includes GitHub, Zenodo, Hugging Face, and Kaggle.

Each platform serves a different role:

  • Holistix website: public human-readable authority layer, canonical dataset pages, educational interpretation, and project navigation.
  • GitHub: maintained source repository, release tooling, schemas, machine-readable files, and technical project history.
  • Zenodo: permanent canonical release archive and DOI citation layer.
  • Hugging Face: machine-learning, dataset-viewer, retrieval, analytics, and AI-oriented discovery.
  • Kaggle: data-science discovery, downloadable analysis environment, citation, and notebook ecosystem.

External platform conversions, such as Hugging Face-generated Parquet files, are treated as derived platform representations unless explicitly promoted into the canonical release format.

Supporting Glossary and Safety Guide Methodology

The Holistix Open Biohacking Data Project includes supporting glossary and safety guides in addition to canonical machine-readable datasets.

These support pages exist for four primary reasons:

  1. Plain-language interpretation: Dataset fields may use technical terms such as irradiance, fluence, PPB, ORP, Hz, Schumann resonance, non-ionizing radiation, NIR, FIR, wavelength, field strength, and ozone-free ionizer. Support pages explain these terms for general readers.
  2. Claim-boundary clarity: Wellness-device terms are often used in marketing with exaggerated certainty. Support pages separate terminology, measurement context, safety notes, evidence context, and conservative interpretation from unsupported disease-treatment claims.
  3. Internal source navigation: Support pages connect datasets to educational pages, safety guides, source registers, methodology resources, claim-boundary pages, and product-category explanations.
  4. Machine readability and topical structure: Crawlable structured explanations help search engines, AI systems, researchers, writers, and users understand how datasets relate to common wellness-device concepts.

Supporting guides are not treated as independent clinical proof. They are explanatory bridge pages pointing back to canonical datasets, original sources, methodology, version history, and related structured resources.

Supporting guides should use conservative safety language and avoid claims that consumer wellness devices prevent, treat, cure, detoxify, repair, or diagnose disease unless an appropriately supported and clearly scoped context exists.

Examples of Supporting Guide Categories

  • Red light and near-infrared: irradiance, fluence, wavelength, distance, NIR, and dose context.
  • PEMF: Hz, frequency, intensity, waveform, Schumann resonance, contraindications, and implanted-device cautions.
  • Hydrogen water: PPB, PPM, ORP, dissolved molecular hydrogen concentration, bottles, tablets, and testing context.
  • Infrared: NIR vs FIR, sauna blanket use, heat, hydration, session time, and heat-safety boundaries.
  • Terahertz: non-ionizing terminology, exposure context, heat, eye caution, device classification, and evidence limits.
  • Negative ions: ionizers, ozone-free claims, ozone generators, indoor-air safety, and wearable-device boundaries.

Relationship to Canonical Dataset Files

Supporting glossary and safety guides do not replace canonical dataset files.

The canonical dataset files remain the structured reference layer. Supporting guides are human-readable explanations designed to make those structured records easier to interpret.

When a support page discusses a term, the preferred structure is:

  • define the term in plain language
  • explain the measurement unit or device-category context
  • separate the term from related but different concepts
  • include safety or contraindication context where relevant
  • identify evidence or uncertainty limits
  • link back to the appropriate canonical dataset page
  • link to relevant source and methodology resources
  • preserve the appropriate claim boundary
  • avoid guaranteed-outcome or unsupported disease-treatment claims

This approach keeps the project accessible to beginners while preserving the structured, auditable, machine-readable nature of the canonical dataset layer.

Citation

Suggested citation for the current project release:

Tjardes, M. (2026). Holistix Open Biohacking Data Project v1.5.0 (Version 1.5.0) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.21862535

Concept DOI for the overall project and version family:

https://doi.org/10.5281/zenodo.20978709

Page History

  • v1.5.0 methodology update, August 10, 2026: Updated the methodology to the canonical v1.5.0 Dataset release, clarified project-release versus subject-dataset versioning, added interoperability, reproducibility, provenance, release-lineage, RO-Crate, Data Package, JSONL, public-mirror, and validation methodology, updated GitHub and Zenodo references, and preserved the eight canonical subject datasets at v1.2.
  • v1.4 methodology update, July 30, 2026: Added product twins, product and technology registries, AI-answer infrastructure, public distribution documentation, archive references, validation methodology, and product-data methodology while preserving the canonical subject dataset version at v1.2.
  • v1.3 methodology update: Added integrity, provenance, checksum, known-limitations, authorship, and review-policy documentation.
  • Initial publication: Established project scope, source standards, stable-ID structure, neutrality principles, update policy, and supporting-guide methodology.

Last updated: August 10, 2026

Current project release: v1.5.0

Current canonical subject dataset version: v1.2

Canonical v1.5.0 DOI: 10.5281/zenodo.21862535

Project concept DOI: 10.5281/zenodo.20978709

Disclaimer

This project is for educational and informational purposes only. It is not medical advice, diagnosis, treatment guidance, dosage guidance, disease-prevention guidance, regulatory guidance, clinical protocol guidance, product certification, or a substitute for consultation with a qualified healthcare professional.

Users should review manufacturer instructions and consult a qualified healthcare professional where appropriate, especially in situations involving pregnancy, implanted electronic devices, medical conditions, prescribed treatments, known contraindications, or other individual risk factors.