From patchwork Mon Jul 25 20:32:03 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: David Marchand X-Patchwork-Id: 114176 X-Patchwork-Delegate: maxime.coquelin@redhat.com Return-Path: X-Original-To: patchwork@inbox.dpdk.org Delivered-To: patchwork@inbox.dpdk.org Received: from mails.dpdk.org (mails.dpdk.org [217.70.189.124]) by inbox.dpdk.org (Postfix) with ESMTP id 39867A00C4; Mon, 25 Jul 2022 22:32:30 +0200 (CEST) Received: from [217.70.189.124] (localhost [127.0.0.1]) by mails.dpdk.org (Postfix) with ESMTP id 173E74280B; Mon, 25 Jul 2022 22:32:27 +0200 (CEST) Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) by mails.dpdk.org (Postfix) with ESMTP id 92E7C41144 for ; Mon, 25 Jul 2022 22:32:25 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1658781145; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Uachlq7yvT8XvDhoD763esWdz1NPhtNbbqDHzHElBg8=; b=bGAqqIDvwHbc2hB8HduACY41FjJo9e19PSvflotXxf2/pMwaXEyu1zKVlxGLN4BiLsVoVY J1cu3nDL7Q3tSrYPfJYR6GMaEbO9l5klDiUrIgLGn97ilzakVcfcOTdbk1XYvgCEIjdffE kn/Umy/+rhIb3anRJKaBDxv2LXLCtvc= Received: from mimecast-mx02.redhat.com (mimecast-mx02.redhat.com [66.187.233.88]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id us-mta-214-o_4UFHiKP4ihkz56UgJNwQ-1; Mon, 25 Jul 2022 16:32:24 -0400 X-MC-Unique: o_4UFHiKP4ihkz56UgJNwQ-1 Received: from smtp.corp.redhat.com (int-mx07.intmail.prod.int.rdu2.redhat.com [10.11.54.7]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by mimecast-mx02.redhat.com (Postfix) with ESMTPS id A984F81F46D; Mon, 25 Jul 2022 20:32:23 +0000 (UTC) Received: from localhost.localdomain (unknown [10.40.192.6]) by smtp.corp.redhat.com (Postfix) with ESMTP id BD6EF141511F; Mon, 25 Jul 2022 20:32:22 +0000 (UTC) From: David Marchand To: dev@dpdk.org Cc: stable@dpdk.org, Maxime Coquelin , Chenbo Xia Subject: [PATCH v3 1/4] vhost: fix vq use after free on NUMA reallocation Date: Mon, 25 Jul 2022 22:32:03 +0200 Message-Id: <20220725203206.427083-2-david.marchand@redhat.com> In-Reply-To: <20220725203206.427083-1-david.marchand@redhat.com> References: <20220722135320.109269-1-david.marchand@redhat.com> <20220725203206.427083-1-david.marchand@redhat.com> MIME-Version: 1.0 X-Scanned-By: MIMEDefang 2.85 on 10.11.54.7 Authentication-Results: relay.mimecast.com; auth=pass smtp.auth=CUSA124A263 smtp.mailfrom=david.marchand@redhat.com X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com X-BeenThere: dev@dpdk.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org translate_ring_addresses (via numa_realloc) may change a virtio device and virtio queue. The virtqueue object must be refreshed before accessing the lock. Fixes: 04c27cb673b9 ("vhost: fix unsafe vring addresses modifications") Cc: stable@dpdk.org Signed-off-by: David Marchand Reviewed-by: Maxime Coquelin --- lib/vhost/vhost_user.c | 1 + 1 file changed, 1 insertion(+) diff --git a/lib/vhost/vhost_user.c b/lib/vhost/vhost_user.c index 4ad28bac45..91d40e32fc 100644 --- a/lib/vhost/vhost_user.c +++ b/lib/vhost/vhost_user.c @@ -2596,6 +2596,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev, if (is_vring_iotlb(dev, vq, imsg)) { rte_spinlock_lock(&vq->access_lock); *pdev = dev = translate_ring_addresses(dev, i); + vq = dev->virtqueue[i]; rte_spinlock_unlock(&vq->access_lock); } } From patchwork Mon Jul 25 20:32:04 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: David Marchand X-Patchwork-Id: 114177 X-Patchwork-Delegate: maxime.coquelin@redhat.com Return-Path: X-Original-To: patchwork@inbox.dpdk.org Delivered-To: patchwork@inbox.dpdk.org Received: from mails.dpdk.org (mails.dpdk.org [217.70.189.124]) by inbox.dpdk.org (Postfix) with ESMTP id 3D30CA00C4; Mon, 25 Jul 2022 22:32:36 +0200 (CEST) Received: from [217.70.189.124] (localhost [127.0.0.1]) by mails.dpdk.org (Postfix) with ESMTP id 2FCB342826; Mon, 25 Jul 2022 22:32:32 +0200 (CEST) Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) by mails.dpdk.org (Postfix) with ESMTP id 44E6E42825 for ; Mon, 25 Jul 2022 22:32:31 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1658781150; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=VheqVruvMbQAUMiykcoyzHQ5JYr10kgpXthD+ZFjJQ8=; b=VbBPtuEuBUarRxeQWS6v9vIoco8G15Xyi/qD3agkKKS8DxVyQ1Qei3LsVypg9WKDS60qcH bN2gDRnW66BQYZN0PvWtiCTilgYjgiMDqOUgXkpmAvtNTO+WpkDj+d0n1UNqnwJOYgDH/t 8Hc9XbaOknUc4+yUAu/ul6Yj94rFyIA= Received: from mimecast-mx02.redhat.com (mimecast-mx02.redhat.com [66.187.233.88]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id us-mta-581-UNIEl5axPq2B3hlY_irYbg-1; Mon, 25 Jul 2022 16:32:27 -0400 X-MC-Unique: UNIEl5axPq2B3hlY_irYbg-1 Received: from smtp.corp.redhat.com (int-mx07.intmail.prod.int.rdu2.redhat.com [10.11.54.7]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by mimecast-mx02.redhat.com (Postfix) with ESMTPS id 4E3AE811E80; Mon, 25 Jul 2022 20:32:27 +0000 (UTC) Received: from localhost.localdomain (unknown [10.40.192.6]) by smtp.corp.redhat.com (Postfix) with ESMTP id 639501415122; Mon, 25 Jul 2022 20:32:26 +0000 (UTC) From: David Marchand To: dev@dpdk.org Cc: Maxime Coquelin , Chenbo Xia Subject: [PATCH v3 2/4] vhost: make NUMA reallocation code more robust Date: Mon, 25 Jul 2022 22:32:04 +0200 Message-Id: <20220725203206.427083-3-david.marchand@redhat.com> In-Reply-To: <20220725203206.427083-1-david.marchand@redhat.com> References: <20220722135320.109269-1-david.marchand@redhat.com> <20220725203206.427083-1-david.marchand@redhat.com> MIME-Version: 1.0 X-Scanned-By: MIMEDefang 2.85 on 10.11.54.7 Authentication-Results: relay.mimecast.com; auth=pass smtp.auth=CUSA124A263 smtp.mailfrom=david.marchand@redhat.com X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com X-BeenThere: dev@dpdk.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org translate_ring_addresses and numa_realloc may change a virtio device and virtio queue. Callers of those helpers must be extra careful and refresh any reference to old data. Change those functions prototype as a way to hint about this issue and always ask for an indirect pointer. Besides, when reallocating the device and queue, the code already made sure it will return a pointer to a valid device. The checks on such returned pointer can be removed. Signed-off-by: David Marchand Reviewed-by: Maxime Coquelin --- lib/vhost/vhost_user.c | 144 +++++++++++++++++++---------------------- 1 file changed, 66 insertions(+), 78 deletions(-) diff --git a/lib/vhost/vhost_user.c b/lib/vhost/vhost_user.c index 91d40e32fc..46d4a02c1e 100644 --- a/lib/vhost/vhost_user.c +++ b/lib/vhost/vhost_user.c @@ -493,11 +493,11 @@ vhost_user_set_vring_num(struct virtio_net **pdev, * make them on the same numa node as the memory of vring descriptor. */ #ifdef RTE_LIBRTE_VHOST_NUMA -static struct virtio_net* -numa_realloc(struct virtio_net *dev, int index) +static void +numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) { int node, dev_node; - struct virtio_net *old_dev; + struct virtio_net *dev; struct vhost_virtqueue *vq; struct batch_copy_elem *bce; struct guest_page *gp; @@ -505,34 +505,35 @@ numa_realloc(struct virtio_net *dev, int index) size_t mem_size; int ret; - old_dev = dev; - vq = dev->virtqueue[index]; + dev = *pdev; + vq = *pvq; /* * If VQ is ready, it is too late to reallocate, it certainly already * happened anyway on VHOST_USER_SET_VRING_ADRR. */ if (vq->ready) - return dev; + return; ret = get_mempolicy(&node, NULL, 0, vq->desc, MPOL_F_NODE | MPOL_F_ADDR); if (ret) { VHOST_LOG_CONFIG(dev->ifname, ERR, "unable to get virtqueue %d numa information.\n", index); - return dev; + return; } if (node == vq->numa_node) goto out_dev_realloc; - vq = rte_realloc_socket(vq, sizeof(*vq), 0, node); + vq = rte_realloc_socket(*pvq, sizeof(**pvq), 0, node); if (!vq) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc virtqueue %d on node %d\n", index, node); - return dev; + return; } + *pvq = vq; if (vq != dev->virtqueue[index]) { VHOST_LOG_CONFIG(dev->ifname, INFO, "reallocated virtqueue on node %d\n", node); @@ -549,7 +550,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc shadow packed on node %d\n", node); - return dev; + return; } vq->shadow_used_packed = sup; } else { @@ -561,7 +562,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc shadow split on node %d\n", node); - return dev; + return; } vq->shadow_used_split = sus; } @@ -572,7 +573,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc batch copy elem on node %d\n", node); - return dev; + return; } vq->batch_copy_elems = bce; @@ -584,7 +585,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc log cache on node %d\n", node); - return dev; + return; } vq->log_cache = lc; } @@ -597,7 +598,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc resubmit inflight on node %d\n", node); - return dev; + return; } vq->resubmit_inflight = ri; @@ -610,7 +611,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc resubmit list on node %d\n", node); - return dev; + return; } ri->resubmit_list = rd; } @@ -621,22 +622,23 @@ numa_realloc(struct virtio_net *dev, int index) out_dev_realloc: if (dev->flags & VIRTIO_DEV_RUNNING) - return dev; + return; ret = get_mempolicy(&dev_node, NULL, 0, dev, MPOL_F_NODE | MPOL_F_ADDR); if (ret) { VHOST_LOG_CONFIG(dev->ifname, ERR, "unable to get numa information.\n"); - return dev; + return; } if (dev_node == node) - return dev; + return; - dev = rte_realloc_socket(old_dev, sizeof(*dev), 0, node); + dev = rte_realloc_socket(*pdev, sizeof(**pdev), 0, node); if (!dev) { - VHOST_LOG_CONFIG(old_dev->ifname, ERR, "failed to realloc dev on node %d\n", node); - return old_dev; + VHOST_LOG_CONFIG((*pdev)->ifname, ERR, "failed to realloc dev on node %d\n", node); + return; } + *pdev = dev; VHOST_LOG_CONFIG(dev->ifname, INFO, "reallocated device on node %d\n", node); vhost_devices[dev->vid] = dev; @@ -648,7 +650,7 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc mem table on node %d\n", node); - return dev; + return; } dev->mem = mem; @@ -658,17 +660,17 @@ numa_realloc(struct virtio_net *dev, int index) VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc guest pages on node %d\n", node); - return dev; + return; } dev->guest_pages = gp; - - return dev; } #else -static struct virtio_net* -numa_realloc(struct virtio_net *dev, int index __rte_unused) +static void +numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) { - return dev; + RTE_SET_USED(pdev); + RTE_SET_USED(pvq); + RTE_SET_USED(index); } #endif @@ -738,88 +740,92 @@ log_addr_to_gpa(struct virtio_net *dev, struct vhost_virtqueue *vq) return log_gpa; } -static struct virtio_net * -translate_ring_addresses(struct virtio_net *dev, int vq_index) +static void +translate_ring_addresses(struct virtio_net **pdev, struct vhost_virtqueue **pvq, + int vq_index) { - struct vhost_virtqueue *vq = dev->virtqueue[vq_index]; - struct vhost_vring_addr *addr = &vq->ring_addrs; + struct vhost_virtqueue *vq; + struct virtio_net *dev; uint64_t len, expected_len; - if (addr->flags & (1 << VHOST_VRING_F_LOG)) { + dev = *pdev; + vq = *pvq; + + if (vq->ring_addrs.flags & (1 << VHOST_VRING_F_LOG)) { vq->log_guest_addr = log_addr_to_gpa(dev, vq); if (vq->log_guest_addr == 0) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to map log_guest_addr.\n"); - return dev; + return; } } if (vq_is_packed(dev)) { len = sizeof(struct vring_packed_desc) * vq->size; vq->desc_packed = (struct vring_packed_desc *)(uintptr_t) - ring_addr_to_vva(dev, vq, addr->desc_user_addr, &len); + ring_addr_to_vva(dev, vq, vq->ring_addrs.desc_user_addr, &len); if (vq->desc_packed == NULL || len != sizeof(struct vring_packed_desc) * vq->size) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to map desc_packed ring.\n"); - return dev; + return; } - dev = numa_realloc(dev, vq_index); - vq = dev->virtqueue[vq_index]; - addr = &vq->ring_addrs; + numa_realloc(&dev, &vq, vq_index); + *pdev = dev; + *pvq = vq; len = sizeof(struct vring_packed_desc_event); vq->driver_event = (struct vring_packed_desc_event *) (uintptr_t)ring_addr_to_vva(dev, - vq, addr->avail_user_addr, &len); + vq, vq->ring_addrs.avail_user_addr, &len); if (vq->driver_event == NULL || len != sizeof(struct vring_packed_desc_event)) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to find driver area address.\n"); - return dev; + return; } len = sizeof(struct vring_packed_desc_event); vq->device_event = (struct vring_packed_desc_event *) (uintptr_t)ring_addr_to_vva(dev, - vq, addr->used_user_addr, &len); + vq, vq->ring_addrs.used_user_addr, &len); if (vq->device_event == NULL || len != sizeof(struct vring_packed_desc_event)) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to find device area address.\n"); - return dev; + return; } vq->access_ok = true; - return dev; + return; } /* The addresses are converted from QEMU virtual to Vhost virtual. */ if (vq->desc && vq->avail && vq->used) - return dev; + return; len = sizeof(struct vring_desc) * vq->size; vq->desc = (struct vring_desc *)(uintptr_t)ring_addr_to_vva(dev, - vq, addr->desc_user_addr, &len); + vq, vq->ring_addrs.desc_user_addr, &len); if (vq->desc == 0 || len != sizeof(struct vring_desc) * vq->size) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to map desc ring.\n"); - return dev; + return; } - dev = numa_realloc(dev, vq_index); - vq = dev->virtqueue[vq_index]; - addr = &vq->ring_addrs; + numa_realloc(&dev, &vq, vq_index); + *pdev = dev; + *pvq = vq; len = sizeof(struct vring_avail) + sizeof(uint16_t) * vq->size; if (dev->features & (1ULL << VIRTIO_RING_F_EVENT_IDX)) len += sizeof(uint16_t); expected_len = len; vq->avail = (struct vring_avail *)(uintptr_t)ring_addr_to_vva(dev, - vq, addr->avail_user_addr, &len); + vq, vq->ring_addrs.avail_user_addr, &len); if (vq->avail == 0 || len != expected_len) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to map avail ring.\n"); - return dev; + return; } len = sizeof(struct vring_used) + @@ -828,10 +834,10 @@ translate_ring_addresses(struct virtio_net *dev, int vq_index) len += sizeof(uint16_t); expected_len = len; vq->used = (struct vring_used *)(uintptr_t)ring_addr_to_vva(dev, - vq, addr->used_user_addr, &len); + vq, vq->ring_addrs.used_user_addr, &len); if (vq->used == 0 || len != expected_len) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, "failed to map used ring.\n"); - return dev; + return; } if (vq->last_used_idx != vq->used->idx) { @@ -850,8 +856,6 @@ translate_ring_addresses(struct virtio_net *dev, int vq_index) VHOST_LOG_CONFIG(dev->ifname, DEBUG, "mapped address avail: %p\n", vq->avail); VHOST_LOG_CONFIG(dev->ifname, DEBUG, "mapped address used: %p\n", vq->used); VHOST_LOG_CONFIG(dev->ifname, DEBUG, "log_guest_addr: %" PRIx64 "\n", vq->log_guest_addr); - - return dev; } /* @@ -887,10 +891,7 @@ vhost_user_set_vring_addr(struct virtio_net **pdev, if ((vq->enabled && (dev->features & (1ULL << VHOST_USER_F_PROTOCOL_FEATURES))) || access_ok) { - dev = translate_ring_addresses(dev, ctx->msg.payload.addr.index); - if (!dev) - return RTE_VHOST_MSG_RESULT_ERR; - + translate_ring_addresses(&dev, &vq, ctx->msg.payload.addr.index); *pdev = dev; } @@ -1396,12 +1397,7 @@ vhost_user_set_mem_table(struct virtio_net **pdev, */ vring_invalidate(dev, vq); - dev = translate_ring_addresses(dev, i); - if (!dev) { - dev = *pdev; - goto free_mem_table; - } - + translate_ring_addresses(&dev, &vq, i); *pdev = dev; } } @@ -2029,17 +2025,9 @@ vhost_user_set_vring_kick(struct virtio_net **pdev, file.index, file.fd); /* Interpret ring addresses only when ring is started. */ - dev = translate_ring_addresses(dev, file.index); - if (!dev) { - if (file.fd != VIRTIO_INVALID_EVENTFD) - close(file.fd); - - return RTE_VHOST_MSG_RESULT_ERR; - } - - *pdev = dev; - vq = dev->virtqueue[file.index]; + translate_ring_addresses(&dev, &vq, file.index); + *pdev = dev; /* * When VHOST_USER_F_PROTOCOL_FEATURES is not negotiated, @@ -2595,8 +2583,8 @@ vhost_user_iotlb_msg(struct virtio_net **pdev, if (is_vring_iotlb(dev, vq, imsg)) { rte_spinlock_lock(&vq->access_lock); - *pdev = dev = translate_ring_addresses(dev, i); - vq = dev->virtqueue[i]; + translate_ring_addresses(&dev, &vq, i); + *pdev = dev; rte_spinlock_unlock(&vq->access_lock); } } From patchwork Mon Jul 25 20:32:05 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: David Marchand X-Patchwork-Id: 114178 X-Patchwork-Delegate: maxime.coquelin@redhat.com Return-Path: X-Original-To: patchwork@inbox.dpdk.org Delivered-To: patchwork@inbox.dpdk.org Received: from mails.dpdk.org (mails.dpdk.org [217.70.189.124]) by inbox.dpdk.org (Postfix) with ESMTP id D64C1A00C4; Mon, 25 Jul 2022 22:32:41 +0200 (CEST) Received: from [217.70.189.124] (localhost [127.0.0.1]) by mails.dpdk.org (Postfix) with ESMTP id 1470042847; Mon, 25 Jul 2022 22:32:34 +0200 (CEST) Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) by mails.dpdk.org (Postfix) with ESMTP id DF0054282D for ; Mon, 25 Jul 2022 22:32:32 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1658781152; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Q9tTMyhd6KegU75vmu+zUidJ2WfW7B5nK6oKe84Mo8Y=; b=Mad/tybXXwRapsca7KrjLI98tE1YMaCk8KCh2y7JOPfZ8MmpVgnfENd2eNQ68+iTFKMgn9 srKR1wAoQmZCV/LKLa6Efe6p98g8yVfOQwHxp2MVAUmwoOL8ehB5GYzZFZNW2lr/hMDIl3 Y8Zr4VkBVdoZm0Ib3uy8sMH03JNnC/8= Received: from mimecast-mx02.redhat.com (mimecast-mx02.redhat.com [66.187.233.88]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id us-mta-386-1QIiTFZZNtSX5xJMaHN5Iw-1; Mon, 25 Jul 2022 16:32:30 -0400 X-MC-Unique: 1QIiTFZZNtSX5xJMaHN5Iw-1 Received: from smtp.corp.redhat.com (int-mx07.intmail.prod.int.rdu2.redhat.com [10.11.54.7]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by mimecast-mx02.redhat.com (Postfix) with ESMTPS id 73076801755; Mon, 25 Jul 2022 20:32:30 +0000 (UTC) Received: from localhost.localdomain (unknown [10.40.192.6]) by smtp.corp.redhat.com (Postfix) with ESMTP id 834FE1415122; Mon, 25 Jul 2022 20:32:29 +0000 (UTC) From: David Marchand To: dev@dpdk.org Cc: Maxime Coquelin , Chenbo Xia Subject: [PATCH v3 3/4] vhost: keep a reference to virtqueue index Date: Mon, 25 Jul 2022 22:32:05 +0200 Message-Id: <20220725203206.427083-4-david.marchand@redhat.com> In-Reply-To: <20220725203206.427083-1-david.marchand@redhat.com> References: <20220722135320.109269-1-david.marchand@redhat.com> <20220725203206.427083-1-david.marchand@redhat.com> MIME-Version: 1.0 X-Scanned-By: MIMEDefang 2.85 on 10.11.54.7 Authentication-Results: relay.mimecast.com; auth=pass smtp.auth=CUSA124A263 smtp.mailfrom=david.marchand@redhat.com X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com X-BeenThere: dev@dpdk.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org Having a back reference to the index of the vq in the dev->virtqueue[] array makes it possible to unify the internal API, with only passing dev and vq. It also allows displaying the vq index in log messages. Remove virtqueue index checks where unneeded (like in static helpers called from a loop on all available virtqueue). Move virtqueue index validity checks the sooner possible. Signed-off-by: David Marchand Reviewed-by: Maxime Coquelin --- Changes since v2: - rebased on top of cleanup that avoids some use-after-free issues in case of allocation failures or numa realloactions, Changes since v1: - fix vq init (that's what happens when only recompiling the vhost library and not relinking testpmd...), --- lib/vhost/iotlb.c | 5 +-- lib/vhost/iotlb.h | 2 +- lib/vhost/vhost.c | 72 +++++++++++++---------------------- lib/vhost/vhost.h | 3 ++ lib/vhost/vhost_user.c | 46 +++++++++++----------- lib/vhost/virtio_net.c | 86 +++++++++++++++++++----------------------- 6 files changed, 91 insertions(+), 123 deletions(-) diff --git a/lib/vhost/iotlb.c b/lib/vhost/iotlb.c index 35b4193606..dd35338ec0 100644 --- a/lib/vhost/iotlb.c +++ b/lib/vhost/iotlb.c @@ -293,10 +293,9 @@ vhost_user_iotlb_flush_all(struct vhost_virtqueue *vq) } int -vhost_user_iotlb_init(struct virtio_net *dev, int vq_index) +vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq) { char pool_name[RTE_MEMPOOL_NAMESIZE]; - struct vhost_virtqueue *vq = dev->virtqueue[vq_index]; int socket = 0; if (vq->iotlb_pool) { @@ -319,7 +318,7 @@ vhost_user_iotlb_init(struct virtio_net *dev, int vq_index) TAILQ_INIT(&vq->iotlb_pending_list); snprintf(pool_name, sizeof(pool_name), "iotlb_%u_%d_%d", - getpid(), dev->vid, vq_index); + getpid(), dev->vid, vq->index); VHOST_LOG_CONFIG(dev->ifname, DEBUG, "IOTLB cache name: %s\n", pool_name); /* If already created, free it and recreate */ diff --git a/lib/vhost/iotlb.h b/lib/vhost/iotlb.h index 8d0ff7473b..738e31e7b9 100644 --- a/lib/vhost/iotlb.h +++ b/lib/vhost/iotlb.h @@ -47,6 +47,6 @@ void vhost_user_iotlb_pending_insert(struct virtio_net *dev, struct vhost_virtqu void vhost_user_iotlb_pending_remove(struct vhost_virtqueue *vq, uint64_t iova, uint64_t size, uint8_t perm); void vhost_user_iotlb_flush_all(struct vhost_virtqueue *vq); -int vhost_user_iotlb_init(struct virtio_net *dev, int vq_index); +int vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq); #endif /* _VHOST_IOTLB_H_ */ diff --git a/lib/vhost/vhost.c b/lib/vhost/vhost.c index 60cb05a0ff..1b17233652 100644 --- a/lib/vhost/vhost.c +++ b/lib/vhost/vhost.c @@ -575,25 +575,14 @@ vring_invalidate(struct virtio_net *dev, struct vhost_virtqueue *vq) } static void -init_vring_queue(struct virtio_net *dev, uint32_t vring_idx) +init_vring_queue(struct virtio_net *dev, struct vhost_virtqueue *vq, + uint32_t vring_idx) { - struct vhost_virtqueue *vq; int numa_node = SOCKET_ID_ANY; - if (vring_idx >= VHOST_MAX_VRING) { - VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to init vring, out of bound (%d)\n", - vring_idx); - return; - } - - vq = dev->virtqueue[vring_idx]; - if (!vq) { - VHOST_LOG_CONFIG(dev->ifname, ERR, "virtqueue not allocated (%d)\n", vring_idx); - return; - } - memset(vq, 0, sizeof(struct vhost_virtqueue)); + vq->index = vring_idx; vq->kickfd = VIRTIO_UNINITIALIZED_EVENTFD; vq->callfd = VIRTIO_UNINITIALIZED_EVENTFD; vq->notif_enable = VIRTIO_UNINITIALIZED_NOTIF; @@ -607,31 +596,16 @@ init_vring_queue(struct virtio_net *dev, uint32_t vring_idx) #endif vq->numa_node = numa_node; - vhost_user_iotlb_init(dev, vring_idx); + vhost_user_iotlb_init(dev, vq); } static void -reset_vring_queue(struct virtio_net *dev, uint32_t vring_idx) +reset_vring_queue(struct virtio_net *dev, struct vhost_virtqueue *vq) { - struct vhost_virtqueue *vq; int callfd; - if (vring_idx >= VHOST_MAX_VRING) { - VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to reset vring, out of bound (%d)\n", - vring_idx); - return; - } - - vq = dev->virtqueue[vring_idx]; - if (!vq) { - VHOST_LOG_CONFIG(dev->ifname, ERR, - "failed to reset vring, virtqueue not allocated (%d)\n", - vring_idx); - return; - } - callfd = vq->callfd; - init_vring_queue(dev, vring_idx); + init_vring_queue(dev, vq, vq->index); vq->callfd = callfd; } @@ -655,7 +629,7 @@ alloc_vring_queue(struct virtio_net *dev, uint32_t vring_idx) } dev->virtqueue[i] = vq; - init_vring_queue(dev, i); + init_vring_queue(dev, vq, i); rte_spinlock_init(&vq->access_lock); vq->avail_wrap_counter = 1; vq->used_wrap_counter = 1; @@ -681,8 +655,16 @@ reset_device(struct virtio_net *dev) dev->protocol_features = 0; dev->flags &= VIRTIO_DEV_BUILTIN_VIRTIO_NET; - for (i = 0; i < dev->nr_vring; i++) - reset_vring_queue(dev, i); + for (i = 0; i < dev->nr_vring; i++) { + struct vhost_virtqueue *vq = dev->virtqueue[i]; + + if (!vq) { + VHOST_LOG_CONFIG(dev->ifname, ERR, + "failed to reset vring, virtqueue not allocated (%d)\n", i); + continue; + } + reset_vring_queue(dev, vq); + } } /* @@ -1661,17 +1643,15 @@ rte_vhost_extern_callback_register(int vid, } static __rte_always_inline int -async_channel_register(int vid, uint16_t queue_id) +async_channel_register(struct virtio_net *dev, struct vhost_virtqueue *vq) { - struct virtio_net *dev = get_device(vid); - struct vhost_virtqueue *vq = dev->virtqueue[queue_id]; struct vhost_async *async; int node = vq->numa_node; if (unlikely(vq->async)) { VHOST_LOG_CONFIG(dev->ifname, ERR, "async register failed: already registered (qid: %d)\n", - queue_id); + vq->index); return -1; } @@ -1679,7 +1659,7 @@ async_channel_register(int vid, uint16_t queue_id) if (!async) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to allocate async metadata (qid: %d)\n", - queue_id); + vq->index); return -1; } @@ -1688,7 +1668,7 @@ async_channel_register(int vid, uint16_t queue_id) if (!async->pkts_info) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to allocate async_pkts_info (qid: %d)\n", - queue_id); + vq->index); goto out_free_async; } @@ -1697,7 +1677,7 @@ async_channel_register(int vid, uint16_t queue_id) if (!async->pkts_cmpl_flag) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to allocate async pkts_cmpl_flag (qid: %d)\n", - queue_id); + vq->index); goto out_free_async; } @@ -1708,7 +1688,7 @@ async_channel_register(int vid, uint16_t queue_id) if (!async->buffers_packed) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to allocate async buffers (qid: %d)\n", - queue_id); + vq->index); goto out_free_inflight; } } else { @@ -1718,7 +1698,7 @@ async_channel_register(int vid, uint16_t queue_id) if (!async->descs_split) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to allocate async descs (qid: %d)\n", - queue_id); + vq->index); goto out_free_inflight; } } @@ -1753,7 +1733,7 @@ rte_vhost_async_channel_register(int vid, uint16_t queue_id) return -1; rte_spinlock_lock(&vq->access_lock); - ret = async_channel_register(vid, queue_id); + ret = async_channel_register(dev, vq); rte_spinlock_unlock(&vq->access_lock); return ret; @@ -1782,7 +1762,7 @@ rte_vhost_async_channel_register_thread_unsafe(int vid, uint16_t queue_id) return -1; } - return async_channel_register(vid, queue_id); + return async_channel_register(dev, vq); } int diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h index 40fac3b7c6..c6260b54cc 100644 --- a/lib/vhost/vhost.h +++ b/lib/vhost/vhost.h @@ -309,6 +309,9 @@ struct vhost_virtqueue { /* Currently unused as polling mode is enabled */ int kickfd; + /* Index of this vq in dev->virtqueue[] */ + uint32_t index; + /* inflight share memory info */ union { struct rte_vhost_inflight_info_split *inflight_split; diff --git a/lib/vhost/vhost_user.c b/lib/vhost/vhost_user.c index 46d4a02c1e..01820902d8 100644 --- a/lib/vhost/vhost_user.c +++ b/lib/vhost/vhost_user.c @@ -240,22 +240,20 @@ vhost_backend_cleanup(struct virtio_net *dev) } static void -vhost_user_notify_queue_state(struct virtio_net *dev, uint16_t index, - int enable) +vhost_user_notify_queue_state(struct virtio_net *dev, struct vhost_virtqueue *vq, + int enable) { struct rte_vdpa_device *vdpa_dev = dev->vdpa_dev; - struct vhost_virtqueue *vq = dev->virtqueue[index]; /* Configure guest notifications on enable */ if (enable && vq->notif_enable != VIRTIO_UNINITIALIZED_NOTIF) vhost_enable_guest_notification(dev, vq, vq->notif_enable); if (vdpa_dev && vdpa_dev->ops->set_vring_state) - vdpa_dev->ops->set_vring_state(dev->vid, index, enable); + vdpa_dev->ops->set_vring_state(dev->vid, vq->index, enable); if (dev->notify_ops->vring_state_changed) - dev->notify_ops->vring_state_changed(dev->vid, - index, enable); + dev->notify_ops->vring_state_changed(dev->vid, vq->index, enable); } /* @@ -494,7 +492,7 @@ vhost_user_set_vring_num(struct virtio_net **pdev, */ #ifdef RTE_LIBRTE_VHOST_NUMA static void -numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) +numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq) { int node, dev_node; struct virtio_net *dev; @@ -519,7 +517,7 @@ numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) if (ret) { VHOST_LOG_CONFIG(dev->ifname, ERR, "unable to get virtqueue %d numa information.\n", - index); + vq->index); return; } @@ -530,15 +528,15 @@ numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) if (!vq) { VHOST_LOG_CONFIG(dev->ifname, ERR, "failed to realloc virtqueue %d on node %d\n", - index, node); + (*pvq)->index, node); return; } *pvq = vq; - if (vq != dev->virtqueue[index]) { + if (vq != dev->virtqueue[vq->index]) { VHOST_LOG_CONFIG(dev->ifname, INFO, "reallocated virtqueue on node %d\n", node); - dev->virtqueue[index] = vq; - vhost_user_iotlb_init(dev, index); + dev->virtqueue[vq->index] = vq; + vhost_user_iotlb_init(dev, vq); } if (vq_is_packed(dev)) { @@ -666,11 +664,10 @@ numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) } #else static void -numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq, int index) +numa_realloc(struct virtio_net **pdev, struct vhost_virtqueue **pvq) { RTE_SET_USED(pdev); RTE_SET_USED(pvq); - RTE_SET_USED(index); } #endif @@ -741,8 +738,7 @@ log_addr_to_gpa(struct virtio_net *dev, struct vhost_virtqueue *vq) } static void -translate_ring_addresses(struct virtio_net **pdev, struct vhost_virtqueue **pvq, - int vq_index) +translate_ring_addresses(struct virtio_net **pdev, struct vhost_virtqueue **pvq) { struct vhost_virtqueue *vq; struct virtio_net *dev; @@ -771,7 +767,7 @@ translate_ring_addresses(struct virtio_net **pdev, struct vhost_virtqueue **pvq, return; } - numa_realloc(&dev, &vq, vq_index); + numa_realloc(&dev, &vq); *pdev = dev; *pvq = vq; @@ -813,7 +809,7 @@ translate_ring_addresses(struct virtio_net **pdev, struct vhost_virtqueue **pvq, return; } - numa_realloc(&dev, &vq, vq_index); + numa_realloc(&dev, &vq); *pdev = dev; *pvq = vq; @@ -891,7 +887,7 @@ vhost_user_set_vring_addr(struct virtio_net **pdev, if ((vq->enabled && (dev->features & (1ULL << VHOST_USER_F_PROTOCOL_FEATURES))) || access_ok) { - translate_ring_addresses(&dev, &vq, ctx->msg.payload.addr.index); + translate_ring_addresses(&dev, &vq); *pdev = dev; } @@ -1397,7 +1393,7 @@ vhost_user_set_mem_table(struct virtio_net **pdev, */ vring_invalidate(dev, vq); - translate_ring_addresses(&dev, &vq, i); + translate_ring_addresses(&dev, &vq); *pdev = dev; } } @@ -1777,7 +1773,7 @@ vhost_user_set_vring_call(struct virtio_net **pdev, if (vq->ready) { vq->ready = false; - vhost_user_notify_queue_state(dev, file.index, 0); + vhost_user_notify_queue_state(dev, vq, 0); } if (vq->callfd >= 0) @@ -2026,7 +2022,7 @@ vhost_user_set_vring_kick(struct virtio_net **pdev, /* Interpret ring addresses only when ring is started. */ vq = dev->virtqueue[file.index]; - translate_ring_addresses(&dev, &vq, file.index); + translate_ring_addresses(&dev, &vq); *pdev = dev; /* @@ -2040,7 +2036,7 @@ vhost_user_set_vring_kick(struct virtio_net **pdev, if (vq->ready) { vq->ready = false; - vhost_user_notify_queue_state(dev, file.index, 0); + vhost_user_notify_queue_state(dev, vq, 0); } if (vq->kickfd >= 0) @@ -2583,7 +2579,7 @@ vhost_user_iotlb_msg(struct virtio_net **pdev, if (is_vring_iotlb(dev, vq, imsg)) { rte_spinlock_lock(&vq->access_lock); - translate_ring_addresses(&dev, &vq, i); + translate_ring_addresses(&dev, &vq); *pdev = dev; rte_spinlock_unlock(&vq->access_lock); } @@ -3148,7 +3144,7 @@ vhost_user_msg_handler(int vid, int fd) if (cur_ready != (vq && vq->ready)) { vq->ready = cur_ready; - vhost_user_notify_queue_state(dev, i, cur_ready); + vhost_user_notify_queue_state(dev, vq, cur_ready); } } diff --git a/lib/vhost/virtio_net.c b/lib/vhost/virtio_net.c index 35fa4670fd..467dfb203f 100644 --- a/lib/vhost/virtio_net.c +++ b/lib/vhost/virtio_net.c @@ -1555,22 +1555,12 @@ virtio_dev_rx_packed(struct virtio_net *dev, } static __rte_always_inline uint32_t -virtio_dev_rx(struct virtio_net *dev, uint16_t queue_id, +virtio_dev_rx(struct virtio_net *dev, struct vhost_virtqueue *vq, struct rte_mbuf **pkts, uint32_t count) { - struct vhost_virtqueue *vq; uint32_t nb_tx = 0; VHOST_LOG_DATA(dev->ifname, DEBUG, "%s\n", __func__); - if (unlikely(!is_valid_virt_queue_idx(queue_id, 0, dev->nr_vring))) { - VHOST_LOG_DATA(dev->ifname, ERR, - "%s: invalid virtqueue idx %d.\n", - __func__, queue_id); - return 0; - } - - vq = dev->virtqueue[queue_id]; - rte_spinlock_lock(&vq->access_lock); if (unlikely(!vq->enabled)) @@ -1620,7 +1610,14 @@ rte_vhost_enqueue_burst(int vid, uint16_t queue_id, return 0; } - return virtio_dev_rx(dev, queue_id, pkts, count); + if (unlikely(!is_valid_virt_queue_idx(queue_id, 0, dev->nr_vring))) { + VHOST_LOG_DATA(dev->ifname, ERR, + "%s: invalid virtqueue idx %d.\n", + __func__, queue_id); + return 0; + } + + return virtio_dev_rx(dev, dev->virtqueue[queue_id], pkts, count); } static __rte_always_inline uint16_t @@ -1669,8 +1666,7 @@ store_dma_desc_info_packed(struct vring_used_elem_packed *s_ring, static __rte_noinline uint32_t virtio_dev_rx_async_submit_split(struct virtio_net *dev, struct vhost_virtqueue *vq, - uint16_t queue_id, struct rte_mbuf **pkts, uint32_t count, - int16_t dma_id, uint16_t vchan_id) + struct rte_mbuf **pkts, uint32_t count, int16_t dma_id, uint16_t vchan_id) { struct buf_vector buf_vec[BUF_VECTOR_MAX]; uint32_t pkt_idx = 0; @@ -1732,7 +1728,7 @@ virtio_dev_rx_async_submit_split(struct virtio_net *dev, struct vhost_virtqueue VHOST_LOG_DATA(dev->ifname, DEBUG, "%s: failed to transfer %u packets for queue %u.\n", - __func__, pkt_err, queue_id); + __func__, pkt_err, vq->index); /* update number of completed packets */ pkt_idx = n_xfer; @@ -1878,8 +1874,7 @@ dma_error_handler_packed(struct vhost_virtqueue *vq, uint16_t slot_idx, static __rte_noinline uint32_t virtio_dev_rx_async_submit_packed(struct virtio_net *dev, struct vhost_virtqueue *vq, - uint16_t queue_id, struct rte_mbuf **pkts, uint32_t count, - int16_t dma_id, uint16_t vchan_id) + struct rte_mbuf **pkts, uint32_t count, int16_t dma_id, uint16_t vchan_id) { uint32_t pkt_idx = 0; uint32_t remained = count; @@ -1924,7 +1919,7 @@ virtio_dev_rx_async_submit_packed(struct virtio_net *dev, struct vhost_virtqueue if (unlikely(pkt_err)) { VHOST_LOG_DATA(dev->ifname, DEBUG, "%s: failed to transfer %u packets for queue %u.\n", - __func__, pkt_err, queue_id); + __func__, pkt_err, vq->index); dma_error_handler_packed(vq, slot_idx, pkt_err, &pkt_idx); } @@ -2045,11 +2040,9 @@ write_back_completed_descs_packed(struct vhost_virtqueue *vq, } static __rte_always_inline uint16_t -vhost_poll_enqueue_completed(struct virtio_net *dev, uint16_t queue_id, - struct rte_mbuf **pkts, uint16_t count, int16_t dma_id, - uint16_t vchan_id) +vhost_poll_enqueue_completed(struct virtio_net *dev, struct vhost_virtqueue *vq, + struct rte_mbuf **pkts, uint16_t count, int16_t dma_id, uint16_t vchan_id) { - struct vhost_virtqueue *vq = dev->virtqueue[queue_id]; struct vhost_async *async = vq->async; struct async_inflight_info *pkts_info = async->pkts_info; uint16_t nr_cpl_pkts = 0; @@ -2156,7 +2149,7 @@ rte_vhost_poll_enqueue_completed(int vid, uint16_t queue_id, goto out; } - n_pkts_cpl = vhost_poll_enqueue_completed(dev, queue_id, pkts, count, dma_id, vchan_id); + n_pkts_cpl = vhost_poll_enqueue_completed(dev, vq, pkts, count, dma_id, vchan_id); vhost_queue_stats_update(dev, vq, pkts, n_pkts_cpl); vq->stats.inflight_completed += n_pkts_cpl; @@ -2216,12 +2209,11 @@ rte_vhost_clear_queue_thread_unsafe(int vid, uint16_t queue_id, } if ((queue_id & 1) == 0) - n_pkts_cpl = vhost_poll_enqueue_completed(dev, queue_id, - pkts, count, dma_id, vchan_id); - else { + n_pkts_cpl = vhost_poll_enqueue_completed(dev, vq, pkts, count, + dma_id, vchan_id); + else n_pkts_cpl = async_poll_dequeue_completed(dev, vq, pkts, count, - dma_id, vchan_id, dev->flags & VIRTIO_DEV_LEGACY_OL_FLAGS); - } + dma_id, vchan_id, dev->flags & VIRTIO_DEV_LEGACY_OL_FLAGS); vhost_queue_stats_update(dev, vq, pkts, n_pkts_cpl); vq->stats.inflight_completed += n_pkts_cpl; @@ -2275,12 +2267,11 @@ rte_vhost_clear_queue(int vid, uint16_t queue_id, struct rte_mbuf **pkts, } if ((queue_id & 1) == 0) - n_pkts_cpl = vhost_poll_enqueue_completed(dev, queue_id, - pkts, count, dma_id, vchan_id); - else { + n_pkts_cpl = vhost_poll_enqueue_completed(dev, vq, pkts, count, + dma_id, vchan_id); + else n_pkts_cpl = async_poll_dequeue_completed(dev, vq, pkts, count, - dma_id, vchan_id, dev->flags & VIRTIO_DEV_LEGACY_OL_FLAGS); - } + dma_id, vchan_id, dev->flags & VIRTIO_DEV_LEGACY_OL_FLAGS); vhost_queue_stats_update(dev, vq, pkts, n_pkts_cpl); vq->stats.inflight_completed += n_pkts_cpl; @@ -2292,19 +2283,12 @@ rte_vhost_clear_queue(int vid, uint16_t queue_id, struct rte_mbuf **pkts, } static __rte_always_inline uint32_t -virtio_dev_rx_async_submit(struct virtio_net *dev, uint16_t queue_id, +virtio_dev_rx_async_submit(struct virtio_net *dev, struct vhost_virtqueue *vq, struct rte_mbuf **pkts, uint32_t count, int16_t dma_id, uint16_t vchan_id) { - struct vhost_virtqueue *vq; uint32_t nb_tx = 0; VHOST_LOG_DATA(dev->ifname, DEBUG, "%s\n", __func__); - if (unlikely(!is_valid_virt_queue_idx(queue_id, 0, dev->nr_vring))) { - VHOST_LOG_DATA(dev->ifname, ERR, - "%s: invalid virtqueue idx %d.\n", - __func__, queue_id); - return 0; - } if (unlikely(!dma_copy_track[dma_id].vchans || !dma_copy_track[dma_id].vchans[vchan_id].pkts_cmpl_flag_addr)) { @@ -2314,8 +2298,6 @@ virtio_dev_rx_async_submit(struct virtio_net *dev, uint16_t queue_id, return 0; } - vq = dev->virtqueue[queue_id]; - rte_spinlock_lock(&vq->access_lock); if (unlikely(!vq->enabled || !vq->async)) @@ -2333,11 +2315,11 @@ virtio_dev_rx_async_submit(struct virtio_net *dev, uint16_t queue_id, goto out; if (vq_is_packed(dev)) - nb_tx = virtio_dev_rx_async_submit_packed(dev, vq, queue_id, - pkts, count, dma_id, vchan_id); + nb_tx = virtio_dev_rx_async_submit_packed(dev, vq, pkts, count, + dma_id, vchan_id); else - nb_tx = virtio_dev_rx_async_submit_split(dev, vq, queue_id, - pkts, count, dma_id, vchan_id); + nb_tx = virtio_dev_rx_async_submit_split(dev, vq, pkts, count, + dma_id, vchan_id); vq->stats.inflight_submitted += nb_tx; @@ -2368,7 +2350,15 @@ rte_vhost_submit_enqueue_burst(int vid, uint16_t queue_id, return 0; } - return virtio_dev_rx_async_submit(dev, queue_id, pkts, count, dma_id, vchan_id); + if (unlikely(!is_valid_virt_queue_idx(queue_id, 0, dev->nr_vring))) { + VHOST_LOG_DATA(dev->ifname, ERR, + "%s: invalid virtqueue idx %d.\n", + __func__, queue_id); + return 0; + } + + return virtio_dev_rx_async_submit(dev, dev->virtqueue[queue_id], pkts, count, + dma_id, vchan_id); } static inline bool From patchwork Mon Jul 25 20:32:06 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: David Marchand X-Patchwork-Id: 114179 X-Patchwork-Delegate: maxime.coquelin@redhat.com Return-Path: X-Original-To: patchwork@inbox.dpdk.org Delivered-To: patchwork@inbox.dpdk.org Received: from mails.dpdk.org (mails.dpdk.org [217.70.189.124]) by inbox.dpdk.org (Postfix) with ESMTP id 76628A00C4; Mon, 25 Jul 2022 22:32:49 +0200 (CEST) Received: from [217.70.189.124] (localhost [127.0.0.1]) by mails.dpdk.org (Postfix) with ESMTP id 587E342825; Mon, 25 Jul 2022 22:32:38 +0200 (CEST) Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) by mails.dpdk.org (Postfix) with ESMTP id 0581042B6F for ; Mon, 25 Jul 2022 22:32:36 +0200 (CEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1658781156; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=mz9BIBC1QcFteTvDagir0tYmgM+7AwiZPccb6t1P2vA=; b=VJ8q+9jzOz64OXw8gc9Dgk6OeNa65U1ntXleI04piLpBTHY5n5zIM4ZHgLMhBvgrrZ5wa3 K1AVQ9mIoy0kdBMid570FFSprFiZfVnR8Zt+D2u4TGOu7ZU3zImVEGDAdIUo29d4Mvumk4 RYUvL1IIoYP+xqPqwjwkRx9Dn3u6e5w= Received: from mimecast-mx02.redhat.com (mimecast-mx02.redhat.com [66.187.233.88]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id us-mta-189-4E3LRxYpMv256TkYSbDrbA-1; Mon, 25 Jul 2022 16:32:34 -0400 X-MC-Unique: 4E3LRxYpMv256TkYSbDrbA-1 Received: from smtp.corp.redhat.com (int-mx07.intmail.prod.int.rdu2.redhat.com [10.11.54.7]) (using TLSv1.2 with cipher AECDH-AES256-SHA (256/256 bits)) (No client certificate requested) by mimecast-mx02.redhat.com (Postfix) with ESMTPS id BDBF18039A6; Mon, 25 Jul 2022 20:32:33 +0000 (UTC) Received: from localhost.localdomain (unknown [10.40.192.6]) by smtp.corp.redhat.com (Postfix) with ESMTP id 010EF1415122; Mon, 25 Jul 2022 20:32:32 +0000 (UTC) From: David Marchand To: dev@dpdk.org Cc: Maxime Coquelin , Chenbo Xia Subject: [PATCH v3 4/4] vhost: stop using mempool for IOTLB cache Date: Mon, 25 Jul 2022 22:32:06 +0200 Message-Id: <20220725203206.427083-5-david.marchand@redhat.com> In-Reply-To: <20220725203206.427083-1-david.marchand@redhat.com> References: <20220722135320.109269-1-david.marchand@redhat.com> <20220725203206.427083-1-david.marchand@redhat.com> MIME-Version: 1.0 X-Scanned-By: MIMEDefang 2.85 on 10.11.54.7 Authentication-Results: relay.mimecast.com; auth=pass smtp.auth=CUSA124A263 smtp.mailfrom=david.marchand@redhat.com X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: redhat.com X-BeenThere: dev@dpdk.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: dev-bounces@dpdk.org A mempool consumes 3 memzones (with the default ring mempool driver). The default DPDK configuration allows RTE_MAX_MEMZONE (2560) memzones. Assuming there is no other memzones that means that we can have a maximum of 853 mempools. In the vhost library, the IOTLB cache code so far was requesting a mempool per vq, which means that at the maximum, the vhost library could request mempools for 426 qps. This limit was recently reached on big systems with a lot of virtio ports (and multiqueue in use). While the limit on mempool count could be something we fix at the DPDK project level, there is no reason to use mempools for the IOTLB cache: - the IOTLB cache entries do not need to be DMA-able and are only used by the current process (in multiprocess context), - getting/putting objects from/in the mempool is always associated with some other locks, so some level of lock contention is already present, We can convert to a malloc'd pool with objects put in a free list protected by a spinlock. Signed-off-by: David Marchand Reviewed-by: Maxime Coquelin --- lib/vhost/iotlb.c | 102 ++++++++++++++++++++++++++++------------------ lib/vhost/iotlb.h | 1 + lib/vhost/vhost.c | 2 +- lib/vhost/vhost.h | 4 +- 4 files changed, 67 insertions(+), 42 deletions(-) diff --git a/lib/vhost/iotlb.c b/lib/vhost/iotlb.c index dd35338ec0..2a78929e78 100644 --- a/lib/vhost/iotlb.c +++ b/lib/vhost/iotlb.c @@ -13,6 +13,7 @@ struct vhost_iotlb_entry { TAILQ_ENTRY(vhost_iotlb_entry) next; + SLIST_ENTRY(vhost_iotlb_entry) next_free; uint64_t iova; uint64_t uaddr; @@ -22,6 +23,28 @@ struct vhost_iotlb_entry { #define IOTLB_CACHE_SIZE 2048 +static struct vhost_iotlb_entry * +vhost_user_iotlb_pool_get(struct vhost_virtqueue *vq) +{ + struct vhost_iotlb_entry *node; + + rte_spinlock_lock(&vq->iotlb_free_lock); + node = SLIST_FIRST(&vq->iotlb_free_list); + if (node != NULL) + SLIST_REMOVE_HEAD(&vq->iotlb_free_list, next_free); + rte_spinlock_unlock(&vq->iotlb_free_lock); + return node; +} + +static void +vhost_user_iotlb_pool_put(struct vhost_virtqueue *vq, + struct vhost_iotlb_entry *node) +{ + rte_spinlock_lock(&vq->iotlb_free_lock); + SLIST_INSERT_HEAD(&vq->iotlb_free_list, node, next_free); + rte_spinlock_unlock(&vq->iotlb_free_lock); +} + static void vhost_user_iotlb_cache_random_evict(struct vhost_virtqueue *vq); @@ -34,7 +57,7 @@ vhost_user_iotlb_pending_remove_all(struct vhost_virtqueue *vq) RTE_TAILQ_FOREACH_SAFE(node, &vq->iotlb_pending_list, next, temp_node) { TAILQ_REMOVE(&vq->iotlb_pending_list, node, next); - rte_mempool_put(vq->iotlb_pool, node); + vhost_user_iotlb_pool_put(vq, node); } rte_rwlock_write_unlock(&vq->iotlb_pending_lock); @@ -66,22 +89,21 @@ vhost_user_iotlb_pending_insert(struct virtio_net *dev, struct vhost_virtqueue * uint64_t iova, uint8_t perm) { struct vhost_iotlb_entry *node; - int ret; - ret = rte_mempool_get(vq->iotlb_pool, (void **)&node); - if (ret) { + node = vhost_user_iotlb_pool_get(vq); + if (node == NULL) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, - "IOTLB pool %s empty, clear entries for pending insertion\n", - vq->iotlb_pool->name); + "IOTLB pool for vq %"PRIu32" empty, clear entries for pending insertion\n", + vq->index); if (!TAILQ_EMPTY(&vq->iotlb_pending_list)) vhost_user_iotlb_pending_remove_all(vq); else vhost_user_iotlb_cache_random_evict(vq); - ret = rte_mempool_get(vq->iotlb_pool, (void **)&node); - if (ret) { + node = vhost_user_iotlb_pool_get(vq); + if (node == NULL) { VHOST_LOG_CONFIG(dev->ifname, ERR, - "IOTLB pool %s still empty, pending insertion failure\n", - vq->iotlb_pool->name); + "IOTLB pool vq %"PRIu32" still empty, pending insertion failure\n", + vq->index); return; } } @@ -113,7 +135,7 @@ vhost_user_iotlb_pending_remove(struct vhost_virtqueue *vq, if ((node->perm & perm) != node->perm) continue; TAILQ_REMOVE(&vq->iotlb_pending_list, node, next); - rte_mempool_put(vq->iotlb_pool, node); + vhost_user_iotlb_pool_put(vq, node); } rte_rwlock_write_unlock(&vq->iotlb_pending_lock); @@ -128,7 +150,7 @@ vhost_user_iotlb_cache_remove_all(struct vhost_virtqueue *vq) RTE_TAILQ_FOREACH_SAFE(node, &vq->iotlb_list, next, temp_node) { TAILQ_REMOVE(&vq->iotlb_list, node, next); - rte_mempool_put(vq->iotlb_pool, node); + vhost_user_iotlb_pool_put(vq, node); } vq->iotlb_cache_nr = 0; @@ -149,7 +171,7 @@ vhost_user_iotlb_cache_random_evict(struct vhost_virtqueue *vq) RTE_TAILQ_FOREACH_SAFE(node, &vq->iotlb_list, next, temp_node) { if (!entry_idx) { TAILQ_REMOVE(&vq->iotlb_list, node, next); - rte_mempool_put(vq->iotlb_pool, node); + vhost_user_iotlb_pool_put(vq, node); vq->iotlb_cache_nr--; break; } @@ -165,22 +187,21 @@ vhost_user_iotlb_cache_insert(struct virtio_net *dev, struct vhost_virtqueue *vq uint64_t size, uint8_t perm) { struct vhost_iotlb_entry *node, *new_node; - int ret; - ret = rte_mempool_get(vq->iotlb_pool, (void **)&new_node); - if (ret) { + new_node = vhost_user_iotlb_pool_get(vq); + if (new_node == NULL) { VHOST_LOG_CONFIG(dev->ifname, DEBUG, - "IOTLB pool %s empty, clear entries for cache insertion\n", - vq->iotlb_pool->name); + "IOTLB pool vq %"PRIu32" empty, clear entries for cache insertion\n", + vq->index); if (!TAILQ_EMPTY(&vq->iotlb_list)) vhost_user_iotlb_cache_random_evict(vq); else vhost_user_iotlb_pending_remove_all(vq); - ret = rte_mempool_get(vq->iotlb_pool, (void **)&new_node); - if (ret) { + new_node = vhost_user_iotlb_pool_get(vq); + if (new_node == NULL) { VHOST_LOG_CONFIG(dev->ifname, ERR, - "IOTLB pool %s still empty, cache insertion failed\n", - vq->iotlb_pool->name); + "IOTLB pool vq %"PRIu32" still empty, cache insertion failed\n", + vq->index); return; } } @@ -198,7 +219,7 @@ vhost_user_iotlb_cache_insert(struct virtio_net *dev, struct vhost_virtqueue *vq * So if iova already in list, assume identical. */ if (node->iova == new_node->iova) { - rte_mempool_put(vq->iotlb_pool, new_node); + vhost_user_iotlb_pool_put(vq, new_node); goto unlock; } else if (node->iova > new_node->iova) { TAILQ_INSERT_BEFORE(node, new_node, next); @@ -235,7 +256,7 @@ vhost_user_iotlb_cache_remove(struct vhost_virtqueue *vq, if (iova < node->iova + node->size) { TAILQ_REMOVE(&vq->iotlb_list, node, next); - rte_mempool_put(vq->iotlb_pool, node); + vhost_user_iotlb_pool_put(vq, node); vq->iotlb_cache_nr--; } } @@ -295,7 +316,7 @@ vhost_user_iotlb_flush_all(struct vhost_virtqueue *vq) int vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq) { - char pool_name[RTE_MEMPOOL_NAMESIZE]; + unsigned int i; int socket = 0; if (vq->iotlb_pool) { @@ -304,6 +325,7 @@ vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq) * just drop all cached and pending entries. */ vhost_user_iotlb_flush_all(vq); + rte_free(vq->iotlb_pool); } #ifdef RTE_LIBRTE_VHOST_NUMA @@ -311,32 +333,32 @@ vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq) socket = 0; #endif + rte_spinlock_init(&vq->iotlb_free_lock); rte_rwlock_init(&vq->iotlb_lock); rte_rwlock_init(&vq->iotlb_pending_lock); + SLIST_INIT(&vq->iotlb_free_list); TAILQ_INIT(&vq->iotlb_list); TAILQ_INIT(&vq->iotlb_pending_list); - snprintf(pool_name, sizeof(pool_name), "iotlb_%u_%d_%d", - getpid(), dev->vid, vq->index); - VHOST_LOG_CONFIG(dev->ifname, DEBUG, "IOTLB cache name: %s\n", pool_name); - - /* If already created, free it and recreate */ - vq->iotlb_pool = rte_mempool_lookup(pool_name); - rte_mempool_free(vq->iotlb_pool); - - vq->iotlb_pool = rte_mempool_create(pool_name, - IOTLB_CACHE_SIZE, sizeof(struct vhost_iotlb_entry), 0, - 0, 0, NULL, NULL, NULL, socket, - RTE_MEMPOOL_F_NO_CACHE_ALIGN | - RTE_MEMPOOL_F_SP_PUT); + vq->iotlb_pool = rte_calloc_socket("iotlb", IOTLB_CACHE_SIZE, + sizeof(struct vhost_iotlb_entry), 0, socket); if (!vq->iotlb_pool) { - VHOST_LOG_CONFIG(dev->ifname, ERR, "Failed to create IOTLB cache pool %s\n", - pool_name); + VHOST_LOG_CONFIG(dev->ifname, ERR, + "Failed to create IOTLB cache pool for vq %"PRIu32"\n", + vq->index); return -1; } + for (i = 0; i < IOTLB_CACHE_SIZE; i++) + vhost_user_iotlb_pool_put(vq, &vq->iotlb_pool[i]); vq->iotlb_cache_nr = 0; return 0; } + +void +vhost_user_iotlb_destroy(struct vhost_virtqueue *vq) +{ + rte_free(vq->iotlb_pool); +} diff --git a/lib/vhost/iotlb.h b/lib/vhost/iotlb.h index 738e31e7b9..e27ebebcf5 100644 --- a/lib/vhost/iotlb.h +++ b/lib/vhost/iotlb.h @@ -48,5 +48,6 @@ void vhost_user_iotlb_pending_remove(struct vhost_virtqueue *vq, uint64_t iova, uint64_t size, uint8_t perm); void vhost_user_iotlb_flush_all(struct vhost_virtqueue *vq); int vhost_user_iotlb_init(struct virtio_net *dev, struct vhost_virtqueue *vq); +void vhost_user_iotlb_destroy(struct vhost_virtqueue *vq); #endif /* _VHOST_IOTLB_H_ */ diff --git a/lib/vhost/vhost.c b/lib/vhost/vhost.c index 1b17233652..aa671f47a3 100644 --- a/lib/vhost/vhost.c +++ b/lib/vhost/vhost.c @@ -394,7 +394,7 @@ free_vq(struct virtio_net *dev, struct vhost_virtqueue *vq) vhost_free_async_mem(vq); rte_free(vq->batch_copy_elems); - rte_mempool_free(vq->iotlb_pool); + vhost_user_iotlb_destroy(vq); rte_free(vq->log_cache); rte_free(vq); } diff --git a/lib/vhost/vhost.h b/lib/vhost/vhost.h index c6260b54cc..782d916ae0 100644 --- a/lib/vhost/vhost.h +++ b/lib/vhost/vhost.h @@ -299,10 +299,12 @@ struct vhost_virtqueue { rte_rwlock_t iotlb_lock; rte_rwlock_t iotlb_pending_lock; - struct rte_mempool *iotlb_pool; + struct vhost_iotlb_entry *iotlb_pool; TAILQ_HEAD(, vhost_iotlb_entry) iotlb_list; TAILQ_HEAD(, vhost_iotlb_entry) iotlb_pending_list; int iotlb_cache_nr; + rte_spinlock_t iotlb_free_lock; + SLIST_HEAD(, vhost_iotlb_entry) iotlb_free_list; /* Used to notify the guest (trigger interrupt) */ int callfd;