Validation

Validators ensure that the MQTT messages on your broker fulfill your requirements.

In schema validators, requirements are specified in schemas that ensure the format of the MQTT message payload.
Validation can include checking for the existence of fields or setting limits for values in the payload.

Configuration

HiveMQ Data Hub data validation is enabled by default. You configure data validation in the <data-validation> section of the <data-hub> element in your HiveMQ config.xml file.

Example default data validation configuration
<hivemq>
    <data-hub>
        <data-validation>
            <enabled>true</enabled>
            <injection-points>
                <injection-point>AFTER_AUTHORIZATION</injection-point>
            </injection-points>
        </data-validation>
    </data-hub>
</hivemq>

Data Validation Injection Point

In HiveMQ 4.29 or higher, you can configure the <injection-points> setting. The injection point defines when Data Hub applies data policies to an incoming PUBLISH message. The injection point determines whether data policies run before or after the HiveMQ extension system. The injection point also determines whether data policies process PUBLISH messages that HiveMQ extensions create. If you do not configure an injection point, HiveMQ uses AFTER_AUTHORIZATION. Only one <injection-point> is supported. If you configure more than one <injection-point>, the HiveMQ configuration is invalid and HiveMQ does not start.

The injection point affects data policies only. Behavior policies always run before the extension system, regardless of the configured injection point.

Message Processing Order

The HiveMQ broker processes a PUBLISH message from an MQTT client in the following order:

  1. The broker decodes the PUBLISH packet.

  2. If the injection point is BEFORE_EXTENSION, Data Hub applies data policies.

  3. The extension system runs all Publish Inbound Interceptors.

  4. The extension system runs the publish authorization.

  5. If the injection point is AFTER_AUTHORIZATION (default) or AFTER_EXTENSION, Data Hub applies data policies.

  6. The broker routes the PUBLISH message to the matching subscribers.

HiveMQ extensions use the Publish Service of the HiveMQ Extension SDK to create PUBLISH messages. These PUBLISH messages do not originate from an MQTT client. The PUBLISH messages from the Publish Service enter the message processing of the broker after the extension system. For example, the HiveMQ Enterprise Extension for Kafka uses the Publish Service to publish messages from Kafka topics to your broker. The HiveMQ Enterprise Extension for Google Cloud Pub/Sub and the HiveMQ Enterprise Extension for Amazon Kinesis publish messages in the same way.
Data Hub applies data policies to PUBLISH messages from the Publish Service only when the injection point is AFTER_EXTENSION.

Available Injection Points

Table 1. Available injection points
Injection point PUBLISH messages from MQTT clients PUBLISH messages from extensions (Publish Service) Additional messages on branches

BEFORE_EXTENSION

Data policies run before the Publish Inbound Interceptors and before authorization.
Interceptors receive the PUBLISH message as modified by your data policies.
Data Hub also processes PUBLISH messages that the client is not authorized to send.

Not processed by data policies.

Not supported.
If a transformation script publishes additional messages on a branch, HiveMQ logs a warning and does not publish the additional messages.
The original message is still processed.

AFTER_AUTHORIZATION (default)

Data policies run after the Publish Inbound Interceptors and after authorization.
Data Hub receives the PUBLISH message as modified by interceptors.
PUBLISH messages that an interceptor prevents or that fail authorization never reach Data Hub.

Not processed by data policies.

Supported.

AFTER_EXTENSION

Same as AFTER_AUTHORIZATION.

Processed by data policies.

Supported.

When you switch the injection point to AFTER_EXTENSION, your data policies also apply to all PUBLISH messages that extensions create through the Publish Service.
Before you switch, review the topic filters and actions of your data policies.
Existing data policies can drop or redirect messages that extensions such as the HiveMQ Enterprise Extension for Kafka publish.
Example configuration to apply data policies to PUBLISH messages from extensions
<hivemq>
    <data-hub>
        <data-validation>
            <enabled>true</enabled>
            <injection-points>
                <injection-point>AFTER_EXTENSION</injection-point>
            </injection-points>
        </data-validation>
    </data-hub>
</hivemq>

Data Policies and PUBLISH Messages from Extensions

When the injection point is AFTER_EXTENSION, data policies process PUBLISH messages from extensions with the following differences:

  • The ${clientId} variable interpolates to the fixed string HiveMQ Extension. The same string appears in the log messages of your policies.

  • The Mqtt.disconnect function cannot disconnect a client, because no client connection exists. Instead, HiveMQ drops the message.

  • The Mqtt.drop function drops the message. If the Dropped Messages Topic add-on is enabled, HiveMQ publishes the dropped message on the $dropped topic.

  • Transformations and the Delivery.redirectTo function work in the same way as for PUBLISH messages from MQTT clients.

Validation in Data Policies

Since MQTT is data agnostic, MQTT clients can publish data to downstream services through the broker regardless of whether the data is valid or not.
In practice, invalid or incorrectly formatted data can cause unpredictable behavior. For example, in a microservice that needs to process sensor data.

The validations section of your HiveMQ Data Hub policy ensures that the MQTT data in your broker is valid, reliable, consistent, and conforms to your predefined standards.

The validations in your policy definition determine how incoming messages are evaluated. Currently, HiveMQ Data Hub data validation supports validators of the type schema only.

  • The array of validators in the validations section lists the validators the policy executes for all incoming MQTT messages.

Each validator can have one of two outcomes:

  • success: All validators evaluate to true.

  • failure: Any validator evaluates to false.

Schema-based Validation in Data Policies

Schema-based data validation is an effective way to enhance the value of your data pipelines.
Validation against appropriately configured schemas can ensure data quality, reduce errors, and improve the overall usability and interoperability of your data.

The HiveMQ Data Hub supports JSON Schema and Protobuf schema validation for your data policies.

To set up schema-based validation in your data policy, set the validator type to schema and define the arguments that you want to use.

  • schemas: Lists an array of one or more schemas that are used for the validation.

    • schemaId: The unique string that references the schema in the HiveMQ Data Hub.

    • version: The version number of the schema to specify a certain version or the latest schema by using "latest".

  • strategy: Defines how the success or failure of the validator is evaluated. Possible entries are ALL-OF and ANY_OF.

    • ALL_OF: Specifies that the validation is only considered successful (success) if all listed schemas are valid, otherwise unsuccessful (failure).

    • ANY_OF: Specifies that the validation is considered successful (success) if any one of the listed schemas is valid, otherwise unsuccessful (failure).

Example minimal validation configuration in a data policy
"validation": {
  "validators": [
    {
      "type": "schema",
      "arguments": {
        "strategy": "ALL_OF",
        "schemas": [
          {
            "schemaId": "gps_coordinates",
            "version": "1"
          }
        ]
      }
    }
  ]
}
To learn more about data validation feature, see our Getting Started with MQTT Data Validation Using HiveMQ Data Hub blog post.