Including fixes from netfilter.

Looks like our attempt to keep the PRs smaller have only prevented
 this one from getting even bigger. In the last 9 days there were
 405 postings explicitly tagged with [PATCH net], vs 687
 with [PATCH net-next]. 37% of posted patches being fixes is pretty
 crazy, and that's likely undercounting because LLM "researchers"
 more often post fixes without knowing to tag the patches for specific
 trees. I don't have historic data.
 
 In any case, we keep adjusting the criteria. The next PR will be smaller.
 
 Current release - regressions:
 
  - net: defer netdev KOBJ_ADD uevent until the device is published,
    previously rtnl_lock would serialize the accesses vs publishing
 
  - net: explicitly cancel work to avoid races with ref tracker exit
 
  - qrtr: ns: raise lookup limit to 128
 
  - eth: hns3: fix speed configuration residue after driver reload
 
 Previous releases - regressions:
 
  - tcp: do not change rcv_ssthresh in tcp_measure_rcv_mss(),
    regressed flows with MSS and scaling_ratio variability
 
  - Revert "net: thunderbolt: Enable end-to-end flow control also
    in transmit", broke some platforms (no packets coming thru)
 
  - eth: stmmac: resume PHY before hardware setup when opening
    the interface
 
 Previous releases - always broken:
 
  - another pile of fixes for less common protocols (SCTP, TLS, SMC etc.)
 
  - close a couple of AF_PACKET bugs and ways it can build skbs
    problematic for the rest of the stack
 
  - bridge: mrp: fix uninitialised bytes on the wire
 
  - net: devmem: prevent net-iov / page mixing, avoid crashes
 
  - eth: atlantic: free RX pages of consumed but not refilled buffers
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmp0z8gACgkQMUZtbf5S
 Irs6QQ//cpHnTe8YpK7XTLak9zsKXep0ObNybwFeGVtO9ZwpIvUJnW0+DQQOto/f
 iaWJd+6kqXo4nOBvJMdMm+xT/xLVVFPscAfnOhi7P4FsHPFecGxP4lsN+Gtn1afL
 bgJ92IUTPfM0LZ5vvOxCIFPsUpNvtm0MNk/AacRKedUJf5JrelkHYKBIz8qNCEOR
 jwdrlrhUMhAozWX1SmPXO9Hx1cKhx5g5CuZ2vDWkca5ofWkOsUb7sXdC/jdMYsFx
 j0JchO8D54Ej5SrO/0z8tojRfPWgmfTlCr3kARu0b70KCV1p2Ep8HnGVGEmMLZGQ
 dvTBB4MzLfCZuakC9yNwSLh4nA1ShOvMj02vxgN61vlFiKKhIFWkeW/EtfGx2E9s
 XStCg+X1FY0r49oKPu7oF7oUQFRP4QGWNpWP1opVEeOsWNRYgu2ZXmvaHD4862K/
 ZylNHnHOu+3Ig+xc+BWFS0T2yi20tGa3LHJgDO3uGwMVlGKKh7tcF0RoyQlalzZg
 RNI8T7u6EJFCaJHTToBK/O1ImroiaBBgTCrxHqEWbP6S7Gkx51UHP7sp2Ggp3n0+
 pYIQGxWogAtkkNHtap4p6WuCmlMacH/CX32Nwl0v0tjEePxoyRSSc8HlbqSX1Ylv
 tJJJJ7t58PS2G4KGhw5H1WGggDuVCcFrwW1bSzHHO+qcz5MoDug=
 =wJxf
 -----END PGP SIGNATURE-----

Merge tag 'net-7.2-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Jakub Kicinski:
 "Including fixes from netfilter.

  Looks like our attempt to keep the PRs smaller have only prevented
  this one from getting even bigger. In the last 9 days there were
  405 postings explicitly tagged with [PATCH net], vs 687 with [PATCH
  net-next]. 37% of posted patches being fixes is pretty crazy, and
  that's likely undercounting because LLM "researchers" more often post
  fixes without knowing to tag the patches for specific trees. I don't
  have historic data.

  In any case, we keep adjusting the criteria. The next PR will be
  smaller.

  Current release - regressions:

   - net: defer netdev KOBJ_ADD uevent until the device is published,
     previously rtnl_lock would serialize the accesses vs publishing

   - net: explicitly cancel work to avoid races with ref tracker exit

   - qrtr: ns: raise lookup limit to 128

   - eth: hns3: fix speed configuration residue after driver reload

  Previous releases - regressions:

   - tcp: do not change rcv_ssthresh in tcp_measure_rcv_mss(), regressed
     flows with MSS and scaling_ratio variability

   - Revert "net: thunderbolt: Enable end-to-end flow control also in
     transmit", broke some platforms (no packets coming thru)

   - eth: stmmac: resume PHY before hardware setup when opening the
     interface

  Previous releases - always broken:

   - another pile of fixes for less common protocols (SCTP, TLS, SMC
     etc.)

   - close a couple of AF_PACKET bugs and ways it can build skbs
     problematic for the rest of the stack

   - bridge: mrp: fix uninitialised bytes on the wire

   - net: devmem: prevent net-iov / page mixing, avoid crashes

   - eth: atlantic: free RX pages of consumed but not refilled buffers"

* tag 'net-7.2-rc7' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (116 commits)
  igc: fix netdev not re-attached after resume if interface is down
  tls: don't abort the connection on signal-interrupted sends
  net: avoid theoretical races with ref drain
  net: Defer netdev KOBJ_ADD uevent until the device is published
  MAINTAINERS: dpll: zl3073x: replace Prathosh Satish with Min Li
  sctp: clear control chunk transport if it is being removed
  net/atm: fix slab-out-of-bounds read in vcc_setsockopt()
  s390/ism: Fix UAF of sba and ieq during ism_dev_exit()
  packet: use consistent hard_header_len in TX_RING send path
  packet: use consistent hard_header_len in non-ring send paths
  net: remove CAP_SYS_RAWIO zero-padding in dev_validate_header
  bnge: Fix resource leak in bnge_init_nic() error path
  ptp: ocp: Fix board ID over-read
  tls: rx: restore msg_iter before TLS 1.3 optimistic retry
  selftests: tls: add a test for splicing onto a full plaintext record
  tls: don't leave a full plaintext sk_msg ring unpushed
  xdp: reject clones that overrun skb_shared_info tailroom
  mptcp: reclaim forward-allocated memory on RX path errors
  mptcp: fastopen: only mark MPTFO subflows with SYN data
  mptcp: pm: fix memory leak from alloc-during-teardown race
  ...
This commit is contained in:
Linus Torvalds 2026-08-06 11:39:20 -07:00
commit 315f4bd234
127 changed files with 1583 additions and 590 deletions

View File

@ -118,12 +118,16 @@ attribute-sets:
doc: >-
The number of seconds after which a keep alive message is sent to the
peer
checks:
max: 86400
-
name: keepalive-timeout
type: u32
doc: >-
The number of seconds from the last activity after which the peer is
assumed dead
checks:
max: 86400
-
name: del-reason
type: u32

View File

@ -11734,6 +11734,7 @@ F: drivers/net/ethernet/hisilicon/hibmcge/
HISILICON NETWORK SUBSYSTEM DRIVER
M: Jian Shen <shenjian15@huawei.com>
M: Jijie Shao <shaojijie@huawei.com>
L: netdev@vger.kernel.org
S: Maintained
W: http://www.hisilicon.com
@ -17837,7 +17838,7 @@ F: drivers/net/wireless/microchip/
MICROCHIP ZL3073X DRIVER
M: Ivan Vecera <ivecera@redhat.com>
M: Prathosh Satish <Prathosh.Satish@microchip.com>
M: Min Li <min.li@microchip.com>
L: netdev@vger.kernel.org
S: Supported
F: Documentation/devicetree/bindings/dpll/microchip,zl30731.yaml

View File

@ -138,6 +138,7 @@ struct dibs_dev *dibs_dev_alloc(void)
dibs = kzalloc_obj(*dibs);
if (!dibs)
return dibs;
spin_lock_init(&dibs->lock);
dibs->dev.release = dibs_dev_release;
dibs->dev.class = &dibs_class;
device_initialize(&dibs->dev);
@ -186,7 +187,6 @@ int dibs_dev_add(struct dibs_dev *dibs)
int i, ret;
max_dmbs = dibs->ops->max_dmbs();
spin_lock_init(&dibs->lock);
dibs->dmb_clientid_arr = kzalloc(max_dmbs, GFP_KERNEL);
if (!dibs->dmb_clientid_arr)
return -ENOMEM;

View File

@ -1534,8 +1534,8 @@ void bond_alb_monitor(struct work_struct *work)
struct bonding *bond = container_of(work, struct bonding,
alb_work.work);
struct alb_bond_info *bond_info = &(BOND_ALB_INFO(bond));
struct slave *slave, *curr;
struct list_head *iter;
struct slave *slave;
if (!bond_has_slaves(bond)) {
atomic_set(&bond_info->tx_rebalance_counter, 0);
@ -1597,9 +1597,11 @@ void bond_alb_monitor(struct work_struct *work)
* because a slave was disabled then
* it can now leave promiscuous mode.
*/
dev_set_promiscuity(rtnl_dereference(bond->curr_active_slave)->dev,
-1);
bond_info->primary_is_promisc = 0;
curr = rtnl_dereference(bond->curr_active_slave);
if (bond_info->primary_is_promisc && curr) {
dev_set_promiscuity(curr->dev, -1);
bond_info->primary_is_promisc = 0;
}
rtnl_unlock();
rcu_read_lock();

View File

@ -171,6 +171,7 @@ struct pdsc {
struct timer_list wdtimer;
unsigned int wdtimer_period;
struct work_struct health_work;
bool health_stopped;
struct devlink_health_reporter *fw_reporter;
u32 fw_recoveries;

View File

@ -470,8 +470,10 @@ static void pdsc_stop_health_thread(struct pdsc *pdsc)
return;
timer_shutdown_sync(&pdsc->wdtimer);
if (pdsc->health_work.func)
cancel_work_sync(&pdsc->health_work);
if (pdsc->health_work.func && !pdsc->health_stopped) {
disable_work_sync(&pdsc->health_work);
pdsc->health_stopped = true;
}
}
static void pdsc_restart_health_thread(struct pdsc *pdsc)
@ -479,6 +481,10 @@ static void pdsc_restart_health_thread(struct pdsc *pdsc)
if (pdsc->pdev->is_virtfn)
return;
if (pdsc->health_stopped) {
enable_work(&pdsc->health_work);
pdsc->health_stopped = false;
}
timer_setup(&pdsc->wdtimer, pdsc_wdtimer_cb, 0);
mod_timer(&pdsc->wdtimer, jiffies + 1);
}
@ -555,7 +561,11 @@ static pci_ers_result_t pdsc_pci_error_detected(struct pci_dev *pdev,
pci_channel_state_t error)
{
if (error == pci_channel_io_frozen) {
struct pdsc *pdsc = pci_get_drvdata(pdev);
pdsc_reset_prepare(pdev);
if (!pdev->is_virtfn)
cancel_work_sync(&pdsc->pci_reset_work);
return PCI_ERS_RESULT_NEED_RESET;
}

View File

@ -360,6 +360,35 @@ bool aq_ring_tx_clean(struct aq_ring_s *self)
return !!budget;
}
void aq_ring_tx_deinit(struct aq_ring_s *self)
{
if (!self)
return;
for (; self->sw_head != self->sw_tail;
self->sw_head = aq_ring_next_dx(self, self->sw_head)) {
struct aq_ring_buff_s *buff = &self->buff_ring[self->sw_head];
struct device *ndev = aq_nic_get_dev(self->aq_nic);
if (buff->is_mapped) {
if (buff->is_sop) {
dma_unmap_single(ndev, buff->pa, buff->len,
DMA_TO_DEVICE);
} else {
dma_unmap_page(ndev, buff->pa, buff->len,
DMA_TO_DEVICE);
}
}
if (buff->is_eop) {
if (buff->skb)
dev_kfree_skb_any(buff->skb);
else if (buff->xdpf)
xdp_return_frame(buff->xdpf);
}
}
}
static void aq_rx_checksum(struct aq_ring_s *self,
struct aq_ring_buff_s *buff,
struct sk_buff *skb)
@ -921,15 +950,29 @@ int aq_ring_rx_fill(struct aq_ring_s *self)
void aq_ring_rx_deinit(struct aq_ring_s *self)
{
if (!self)
unsigned int i;
if (!self || !self->buff_ring)
return;
for (; self->sw_head != self->sw_tail;
self->sw_head = aq_ring_next_dx(self, self->sw_head)) {
struct aq_ring_buff_s *buff = &self->buff_ring[self->sw_head];
/* Release every page still owned by the ring.
*
* Walking [sw_head, sw_tail) is not enough: refill is batched
* (aq_ring_rx_fill() waits for AQ_CFG_RX_REFILL_THRES free slots),
* so slots that were cleaned but not yet reposted accumulate in the
* [sw_tail, sw_head) gap, and they keep their page for reuse. Walk
* the whole ring and release whatever is left.
*/
for (i = 0; i < self->size; i++) {
struct aq_ring_buff_s *buff = &self->buff_ring[i];
if (!buff->rxdata.page)
continue;
aq_free_rxpage(&buff->rxdata, aq_nic_get_dev(self->aq_nic));
}
self->sw_head = self->sw_tail;
}
void aq_ring_free(struct aq_ring_s *self)

View File

@ -202,6 +202,7 @@ void aq_ring_update_queue_state(struct aq_ring_s *ring);
void aq_ring_queue_wake(struct aq_ring_s *ring);
void aq_ring_queue_stop(struct aq_ring_s *ring);
bool aq_ring_tx_clean(struct aq_ring_s *self);
void aq_ring_tx_deinit(struct aq_ring_s *self);
int aq_xdp_xmit(struct net_device *dev, int num_frames,
struct xdp_frame **frames, u32 flags);
int aq_ring_rx_clean(struct aq_ring_s *self,

View File

@ -275,7 +275,7 @@ void aq_vec_deinit(struct aq_vec_s *self)
for (i = 0U; self->tx_rings > i; ++i) {
ring = self->ring[i];
aq_ring_tx_clean(&ring[AQ_VEC_TX_ID]);
aq_ring_tx_deinit(&ring[AQ_VEC_TX_ID]);
aq_ring_rx_deinit(&ring[AQ_VEC_RX_ID]);
}

View File

@ -141,12 +141,15 @@ static void bnge_aux_dev_release(struct device *dev)
{
struct bnge_auxr_priv *aux_priv =
container_of(dev, struct bnge_auxr_priv, aux_dev.dev);
struct bnge_dev *bd = pci_get_drvdata(aux_priv->auxr_dev->pdev);
struct bnge_auxr_dev *auxr_dev = aux_priv->auxr_dev;
struct bnge_dev *bd = pci_get_drvdata(to_pci_dev(dev->parent));
ida_free(&bnge_aux_dev_ids, aux_priv->id);
kfree(aux_priv->auxr_dev->auxr_info);
if (auxr_dev) {
kfree(auxr_dev->auxr_info);
kfree(auxr_dev);
}
bd->auxr_dev = NULL;
kfree(aux_priv->auxr_dev);
kfree(aux_priv);
bd->aux_priv = NULL;
}

View File

@ -2768,8 +2768,6 @@ static int bnge_init_nic(struct bnge_net *bn)
err_free_ring_grps:
bnge_free_ring_grps(bn);
return rc;
err_free_rx_ring_pair_bufs:
bnge_free_rx_ring_pair_bufs(bn);
return rc;

View File

@ -163,7 +163,8 @@ static int bnge_adjust_rings(struct bnge_dev *bd, u16 *rx,
u16 tx_chunks = bnge_num_tx_to_cp(bd, *tx);
if (tx_chunks != *tx) {
u16 tx_saved = tx_chunks, rc;
u16 tx_saved = tx_chunks;
int rc;
rc = bnge_fix_rings_count(rx, &tx_chunks, max_nq, sh);
if (rc)

View File

@ -4611,11 +4611,14 @@ static void bnxt_init_one_rx_agg_ring_rxbd(struct bnxt *bp,
type = ((u32)rxr->rx_page_size << RX_BD_LEN_SHIFT) |
RX_BD_TYPE_RX_AGG_BD;
/* On P7, setting EOP will cause the chip to disable
* Relaxed Ordering (RO) for TPA data. Disable EOP for
* potentially higher performance with RO.
/* Disable EOP if TPA is enabled to prevent overlapping zero
* padding with the next segment's data. On P7_PLUS, EOP will
* automatically disable Relaxed Ordering (RO) to prevent
* potential data corruption (and may degrade performance). On
* older chips, RO will not be automatically disabled and may
* cause corruption.
*/
if (BNXT_CHIP_P5_AND_MINUS(bp) || !(bp->flags & BNXT_FLAG_TPA))
if (!(bp->flags & BNXT_FLAG_TPA))
type |= RX_BD_FLAGS_AGG_EOP;
bnxt_init_rxbd_pages(ring, type);
@ -6704,22 +6707,36 @@ int bnxt_get_nr_rss_ctxs(struct bnxt *bp, int rx_rings)
static void bnxt_fill_hw_rss_tbl(struct bnxt *bp, struct bnxt_vnic_info *vnic)
{
bool no_rss = !(vnic->flags & BNXT_VNIC_RSS_FLAG);
u16 i, j;
u16 i, j, min_j = bp->rx_nr_rings - 1;
if (!vnic->rss_table)
goto skip_rss_tbl;
/* Fill the RSS indirection table with ring group ids */
for (i = 0, j = 0; i < HW_HASH_INDEX_SIZE; i++) {
if (!no_rss)
j = bp->rss_indir_tbl[i];
min_j = min(j, min_j);
vnic->rss_table[i] = cpu_to_le16(vnic->fw_grp_ids[j]);
}
skip_rss_tbl:
if (vnic->rss_table && !no_rss)
vnic->default_rx_ring = min_j;
else if (vnic->flags & BNXT_VNIC_RFS_FLAG)
vnic->default_rx_ring = vnic->vnic_id - 1;
else if ((vnic->vnic_id == 1) && BNXT_CHIP_TYPE_NITRO_A0(bp))
vnic->default_rx_ring = bp->rx_nr_rings - 1;
else
vnic->default_rx_ring = 0;
}
static void bnxt_fill_hw_rss_tbl_p5(struct bnxt *bp,
struct bnxt_vnic_info *vnic)
{
u16 tbl_size, i, min_j = bp->rx_nr_rings - 1;
__le16 *ring_tbl = vnic->rss_table;
struct bnxt_rx_ring_info *rxr;
u16 tbl_size, i;
tbl_size = bnxt_get_rxfh_indir_size(bp->dev);
@ -6732,6 +6749,7 @@ static void bnxt_fill_hw_rss_tbl_p5(struct bnxt *bp,
j = ethtool_rxfh_context_indir(vnic->rss_ctx)[i];
else
j = bp->rss_indir_tbl[i];
min_j = min(j, min_j);
rxr = &bp->rx_ring[j];
ring_id = rxr->rx_ring_struct.fw_ring_id;
@ -6739,19 +6757,15 @@ static void bnxt_fill_hw_rss_tbl_p5(struct bnxt *bp,
ring_id = bnxt_cp_ring_for_rx(bp, rxr);
*ring_tbl++ = cpu_to_le16(ring_id);
}
vnic->default_rx_ring = min_j;
}
static void
__bnxt_hwrm_vnic_set_rss(struct bnxt *bp, struct hwrm_vnic_rss_cfg_input *req,
struct bnxt_vnic_info *vnic)
{
if (bp->flags & BNXT_FLAG_CHIP_P5_PLUS) {
bnxt_fill_hw_rss_tbl_p5(bp, vnic);
if (bp->flags & BNXT_FLAG_CHIP_P7)
req->flags |= VNIC_RSS_CFG_REQ_FLAGS_IPSEC_HASH_TYPE_CFG_SUPPORT;
} else {
bnxt_fill_hw_rss_tbl(bp, vnic);
}
if (bp->flags & BNXT_FLAG_CHIP_P7)
req->flags |= VNIC_RSS_CFG_REQ_FLAGS_IPSEC_HASH_TYPE_CFG_SUPPORT;
if (bp->rss_hash_delta) {
req->hash_type = cpu_to_le32(bp->rss_hash_delta);
@ -6803,6 +6817,7 @@ static int bnxt_hwrm_vnic_set_rss_p5(struct bnxt *bp,
if (!set_rss)
return hwrm_req_send(bp, req);
bnxt_fill_hw_rss_tbl_p5(bp, vnic);
__bnxt_hwrm_vnic_set_rss(bp, req, vnic);
ring_tbl_map = vnic->rss_table_dma_addr;
nr_ctxs = bnxt_get_nr_rss_ctxs(bp, bp->rx_nr_rings);
@ -6939,8 +6954,9 @@ int bnxt_hwrm_vnic_cfg(struct bnxt *bp, struct bnxt_vnic_info *vnic)
return rc;
if (bp->flags & BNXT_FLAG_CHIP_P5_PLUS) {
struct bnxt_rx_ring_info *rxr = &bp->rx_ring[0];
struct bnxt_rx_ring_info *rxr;
rxr = &bp->rx_ring[vnic->default_rx_ring];
req->default_rx_ring_id =
cpu_to_le16(rxr->rx_ring_struct.fw_ring_id);
req->default_cmpl_ring_id =
@ -6973,13 +6989,7 @@ int bnxt_hwrm_vnic_cfg(struct bnxt *bp, struct bnxt_vnic_info *vnic)
req->cos_rule = cpu_to_le16(0xffff);
}
if (vnic->flags & BNXT_VNIC_RSS_FLAG)
ring = 0;
else if (vnic->flags & BNXT_VNIC_RFS_FLAG)
ring = vnic->vnic_id - 1;
else if ((vnic->vnic_id == 1) && BNXT_CHIP_TYPE_NITRO_A0(bp))
ring = bp->rx_nr_rings - 1;
ring = vnic->default_rx_ring;
grp_idx = bp->rx_ring[ring].bnapi->index;
req->dflt_ring_grp = cpu_to_le16(bp->grp_info[grp_idx].fw_grp_id);
req->lb_rule = cpu_to_le16(0xffff);
@ -10866,6 +10876,7 @@ static int __bnxt_setup_vnic(struct bnxt *bp, struct bnxt_vnic_info *vnic)
}
skip_rss_ctx:
bnxt_fill_hw_rss_tbl(bp, vnic);
/* configure default vnic, ring grp */
rc = bnxt_hwrm_vnic_cfg(bp, vnic);
if (rc) {
@ -11090,6 +11101,11 @@ static int bnxt_set_vnic_mru_p5(struct bnxt *bp, struct bnxt_vnic_info *vnic,
vnic->vnic_id, rc);
return rc;
}
if (rxr_id == vnic->default_rx_ring) {
rc = bnxt_hwrm_vnic_cfg(bp, vnic);
if (rc)
return rc;
}
}
vnic->mru = mru;
bnxt_hwrm_vnic_update(bp, vnic,
@ -11171,6 +11187,9 @@ static int bnxt_setup_nitroa0_vnic(struct bnxt *bp)
return rc;
}
/* Setup the proper default RX ring */
bnxt_fill_hw_rss_tbl(bp, vnic);
rc = bnxt_hwrm_vnic_cfg(bp, vnic);
if (rc) {
netdev_err(bp->dev, "Cannot allocate special vnic for NS2 A0: %x\n",
@ -16217,6 +16236,7 @@ static int bnxt_queue_mem_alloc(struct net_device *dev,
clone->rx_next_cons = 0;
clone->need_head_pool = false;
clone->rx_page_size = qcfg->rx_page_size;
clone->rx_agg_bmap = NULL;
rc = bnxt_alloc_rx_page_pool(bp, clone, rxr->page_pool->p.nid);
if (rc)
@ -16269,6 +16289,8 @@ static int bnxt_queue_mem_alloc(struct net_device *dev,
bnxt_free_one_tpa_info(bp, clone);
err_free_rx_agg_ring:
bnxt_free_ring(bp, &clone->rx_agg_ring_struct.ring_mem);
kfree(clone->rx_agg_bmap);
clone->rx_agg_bmap = NULL;
err_free_rx_ring:
bnxt_free_ring(bp, &clone->rx_ring_struct.ring_mem);
err_rxq_info_unreg:

View File

@ -1334,6 +1334,7 @@ struct bnxt_vnic_info {
#define BNXT_VNIC_RSSCTX_FLAG 0x40
struct ethtool_rxfh_context *rss_ctx;
u32 vnic_id;
u16 default_rx_ring;
};
struct bnxt_rss_ctx {

View File

@ -495,12 +495,15 @@ static int bnxt_ptp_enable(struct ptp_clock_info *ptp_info,
return rc;
case PTP_CLK_REQ_PPS:
/* Configure PHC PPS IN */
rc = bnxt_ptp_cfg_pin(bp, 0, BNXT_PPS_PIN_PPS_IN);
pin_id = 0;
if (!on)
break;
rc = bnxt_ptp_cfg_pin(bp, pin_id, BNXT_PPS_PIN_PPS_IN);
if (rc)
return rc;
rc = bnxt_ptp_cfg_event(bp, BNXT_PPS_EVENT_INTERNAL);
if (!rc)
ptp->pps_info.pins[0].event = BNXT_PPS_EVENT_INTERNAL;
ptp->pps_info.pins[pin_id].event = BNXT_PPS_EVENT_INTERNAL;
return rc;
default:
netdev_err(ptp->bp->dev, "Unrecognized PIN function\n");

View File

@ -3011,8 +3011,9 @@ static void enic_remove(struct pci_dev *pdev)
if (netdev) {
struct enic *enic = netdev_priv(netdev);
cancel_work_sync(&enic->reset);
cancel_work_sync(&enic->change_mtu_work);
disable_work_sync(&enic->reset);
disable_work_sync(&enic->tx_hang_reset);
disable_work_sync(&enic->change_mtu_work);
unregister_netdev(netdev);
enic_dev_deinit(enic);
vnic_dev_close(enic->vdev);

View File

@ -1282,7 +1282,6 @@ static void hix5hd2_dev_remove(struct platform_device *pdev)
struct net_device *ndev = platform_get_drvdata(pdev);
struct hix5hd2_priv *priv = netdev_priv(ndev);
netif_napi_del(&priv->napi);
unregister_netdev(ndev);
mdiobus_unregister(priv->bus);
mdiobus_free(priv->bus);

View File

@ -9498,12 +9498,8 @@ static int hclge_init_ae_dev(struct hnae3_ae_dev *ae_dev)
if (ret)
goto err_ptp_uninit;
if (hdev->hw.mac.media_type != HNAE3_MEDIA_TYPE_COPPER) {
if (hdev->hw.mac.media_type != HNAE3_MEDIA_TYPE_COPPER)
hdev->hw.mac.req_autoneg = hdev->hw.mac.autoneg;
if (hdev->hw.mac.autoneg == AUTONEG_DISABLE &&
hdev->hw.mac.speed != SPEED_UNKNOWN)
hdev->hw.mac.req_speed = hdev->hw.mac.speed;
}
ret = hclge_set_autoneg_speed_dup(hdev);
if (ret) {

View File

@ -3082,7 +3082,7 @@ static void igc_xdp_xmit_zc(struct igc_ring *ring)
meta_req.tx_buffer = bi;
meta_req.meta = meta;
meta_req.used_desc = 0;
xsk_tx_metadata_request(meta, &igc_xsk_tx_metadata_ops,
xsk_tx_metadata_request(pool, &meta, &igc_xsk_tx_metadata_ops,
&meta_req);
/* xsk_tx_metadata_request() may have updated next_to_use */
@ -7585,11 +7585,13 @@ static int __igc_resume(struct device *dev, bool rpm)
err = __igc_open(netdev, true);
if (!rpm)
rtnl_unlock();
if (!err)
netif_device_attach(netdev);
if (err)
return err;
}
return err;
netif_device_attach(netdev);
return 0;
}
static int igc_resume(struct device *dev)

View File

@ -54,10 +54,12 @@ static void otx2_get_egress_burst_cfg(struct otx2_nic *nic, u32 burst,
if (burst) {
*burst_exp = ilog2(burst) ? ilog2(burst) - 1 : 0;
tmp = burst - rounddown_pow_of_two(burst);
if (burst < max_mantissa)
if (burst <= max_mantissa) {
*burst_mantissa = tmp * 2;
else
} else {
WARN_ON(*burst_exp < 7);
*burst_mantissa = tmp / (1ULL << (*burst_exp - 7));
}
} else {
*burst_exp = MAX_BURST_EXPONENT;
*burst_mantissa = max_mantissa;

View File

@ -684,6 +684,9 @@ static int prestera_fw_hdr_parse(struct prestera_fw *fw)
struct prestera_fw_header *hdr;
u32 magic;
if (fw->bin->size < sizeof(*hdr))
return -EINVAL;
hdr = (struct prestera_fw_header *)fw->bin->data;
magic = be32_to_cpu(hdr->magic_number);

View File

@ -1025,13 +1025,11 @@ struct mlx5_fw_tracer *mlx5_fw_tracer_create(struct mlx5_core_dev *dev)
tracer = kvzalloc_obj(*tracer);
if (!tracer)
return ERR_PTR(-ENOMEM);
return NULL;
tracer->work_queue = create_singlethread_workqueue("mlx5_fw_tracer");
if (!tracer->work_queue) {
err = -ENOMEM;
if (!tracer->work_queue)
goto free_tracer;
}
tracer->dev = dev;
@ -1073,7 +1071,7 @@ struct mlx5_fw_tracer *mlx5_fw_tracer_create(struct mlx5_core_dev *dev)
destroy_workqueue(tracer->work_queue);
free_tracer:
kvfree(tracer);
return ERR_PTR(err);
return NULL;
}
static int fw_tracer_event(struct notifier_block *nb, unsigned long action, void *data);
@ -1084,7 +1082,7 @@ int mlx5_fw_tracer_init(struct mlx5_fw_tracer *tracer)
struct mlx5_core_dev *dev;
int err;
if (IS_ERR_OR_NULL(tracer))
if (!tracer)
return 0;
if (!tracer->str_db.loaded)
@ -1134,7 +1132,7 @@ int mlx5_fw_tracer_init(struct mlx5_fw_tracer *tracer)
/* Stop tracer + Cleanup HW resources */
void mlx5_fw_tracer_cleanup(struct mlx5_fw_tracer *tracer)
{
if (IS_ERR_OR_NULL(tracer))
if (!tracer)
return;
mutex_lock(&tracer->state_lock);
@ -1163,7 +1161,7 @@ void mlx5_fw_tracer_cleanup(struct mlx5_fw_tracer *tracer)
/* Free software resources (Buffers, etc ..) */
void mlx5_fw_tracer_destroy(struct mlx5_fw_tracer *tracer)
{
if (IS_ERR_OR_NULL(tracer))
if (!tracer)
return;
mlx5_core_dbg(tracer->dev, "FWTracer: Destroy\n");
@ -1215,7 +1213,7 @@ int mlx5_fw_tracer_reload(struct mlx5_fw_tracer *tracer)
struct mlx5_core_dev *dev;
int err;
if (IS_ERR_OR_NULL(tracer))
if (!tracer)
return 0;
dev = tracer->dev;

View File

@ -483,7 +483,7 @@ typedef int (*mlx5e_fp_xmit_xdp_frame_check)(struct mlx5e_xdpsq *);
typedef bool (*mlx5e_fp_xmit_xdp_frame)(struct mlx5e_xdpsq *,
struct mlx5e_xmit_data *,
int,
struct xsk_tx_metadata *);
struct xsk_tx_metadata **);
struct mlx5e_xdpsq {
/* data path */

View File

@ -30,6 +30,7 @@ enum {
MLX5E_TC_FLOW_FLAG_FAILED = MLX5E_TC_FLOW_BASE + 9,
MLX5E_TC_FLOW_FLAG_SAMPLE = MLX5E_TC_FLOW_BASE + 10,
MLX5E_TC_FLOW_FLAG_USE_ACT_STATS = MLX5E_TC_FLOW_BASE + 11,
MLX5E_TC_FLOW_FLAG_PEER = MLX5E_TC_FLOW_BASE + 12,
};
struct mlx5e_tc_flow_parse_attr {

View File

@ -452,11 +452,11 @@ INDIRECT_CALLABLE_SCOPE int mlx5e_xmit_xdp_frame_check_mpwqe(struct mlx5e_xdpsq
INDIRECT_CALLABLE_SCOPE bool
mlx5e_xmit_xdp_frame(struct mlx5e_xdpsq *sq, struct mlx5e_xmit_data *xdptxd,
int check_result, struct xsk_tx_metadata *meta);
int check_result, struct xsk_tx_metadata **meta);
INDIRECT_CALLABLE_SCOPE bool
mlx5e_xmit_xdp_frame_mpwqe(struct mlx5e_xdpsq *sq, struct mlx5e_xmit_data *xdptxd,
int check_result, struct xsk_tx_metadata *meta)
int check_result, struct xsk_tx_metadata **meta)
{
struct mlx5e_tx_mpwqe *session = &sq->mpwqe;
struct mlx5e_xdpsq_stats *stats = sq->stats;
@ -504,7 +504,10 @@ mlx5e_xmit_xdp_frame_mpwqe(struct mlx5e_xdpsq *sq, struct mlx5e_xmit_data *xdptx
* and it's safe to complete it at any time.
*/
mlx5e_xdp_mpwqe_session_start(sq);
xsk_tx_metadata_request(meta, &mlx5e_xsk_tx_metadata_ops, &session->wqe->eth);
if (meta)
xsk_tx_metadata_request(sq->xsk_pool, meta,
&mlx5e_xsk_tx_metadata_ops,
&session->wqe->eth);
}
mlx5e_xdp_mpwqe_add_dseg(sq, p, stats);
@ -535,7 +538,7 @@ INDIRECT_CALLABLE_SCOPE int mlx5e_xmit_xdp_frame_check(struct mlx5e_xdpsq *sq)
INDIRECT_CALLABLE_SCOPE bool
mlx5e_xmit_xdp_frame(struct mlx5e_xdpsq *sq, struct mlx5e_xmit_data *xdptxd,
int check_result, struct xsk_tx_metadata *meta)
int check_result, struct xsk_tx_metadata **meta)
{
struct mlx5e_xmit_data_frags *xdptxdf =
container_of(xdptxd, struct mlx5e_xmit_data_frags, xd);
@ -649,7 +652,9 @@ mlx5e_xmit_xdp_frame(struct mlx5e_xdpsq *sq, struct mlx5e_xmit_data *xdptxd,
sq->pc += num_wqebbs;
xsk_tx_metadata_request(meta, &mlx5e_xsk_tx_metadata_ops, eseg);
if (meta)
xsk_tx_metadata_request(sq->xsk_pool, meta,
&mlx5e_xsk_tx_metadata_ops, eseg);
sq->doorbell_cseg = cseg;

View File

@ -114,11 +114,11 @@ extern const struct xsk_tx_metadata_ops mlx5e_xsk_tx_metadata_ops;
INDIRECT_CALLABLE_DECLARE(bool mlx5e_xmit_xdp_frame_mpwqe(struct mlx5e_xdpsq *sq,
struct mlx5e_xmit_data *xdptxd,
int check_result,
struct xsk_tx_metadata *meta));
struct xsk_tx_metadata **meta));
INDIRECT_CALLABLE_DECLARE(bool mlx5e_xmit_xdp_frame(struct mlx5e_xdpsq *sq,
struct mlx5e_xmit_data *xdptxd,
int check_result,
struct xsk_tx_metadata *meta));
struct xsk_tx_metadata **meta));
INDIRECT_CALLABLE_DECLARE(int mlx5e_xmit_xdp_frame_check_mpwqe(struct mlx5e_xdpsq *sq));
INDIRECT_CALLABLE_DECLARE(int mlx5e_xmit_xdp_frame_check(struct mlx5e_xdpsq *sq));

View File

@ -105,7 +105,7 @@ bool mlx5e_xsk_tx(struct mlx5e_xdpsq *sq, unsigned int budget)
ret = INDIRECT_CALL_2(sq->xmit_xdp_frame, mlx5e_xmit_xdp_frame_mpwqe,
mlx5e_xmit_xdp_frame, sq, &xdptxd,
check_result, meta);
check_result, &meta);
if (unlikely(!ret)) {
if (sq->mpwqe.wqe)
mlx5e_xdp_mpwqe_complete(sq);

View File

@ -1939,8 +1939,10 @@ int mlx5e_open_txqsq(struct mlx5e_channel *c, u32 tisn, int txq_ix,
void mlx5e_activate_txqsq(struct mlx5e_txqsq *sq)
{
sq->txq = netdev_get_tx_queue(sq->netdev, sq->txq_ix);
/* Reset BQL only when the SQ has no bytes in flight. */
if (sq->cc == sq->pc)
netdev_tx_reset_queue(sq->txq);
set_bit(MLX5E_SQ_STATE_ENABLED, &sq->state);
netdev_tx_reset_queue(sq->txq);
netif_tx_start_queue(sq->txq);
netif_queue_set_napi(sq->netdev, sq->txq_ix, NETDEV_QUEUE_TYPE_TX, sq->cq.napi);
}

View File

@ -2161,7 +2161,8 @@ static void mlx5e_tc_del_flow(struct mlx5e_priv *priv,
if (mlx5e_is_eswitch_flow(flow)) {
struct mlx5_devcom_comp_dev *devcom = flow->priv->mdev->priv.eswitch->devcom;
if (!mlx5_devcom_for_each_peer_begin(devcom)) {
if (flow_flag_test(flow, PEER) ||
!mlx5_devcom_for_each_peer_begin(devcom)) {
mlx5e_tc_del_fdb_flow(priv, flow);
return;
}
@ -4628,6 +4629,7 @@ static int mlx5e_tc_add_fdb_peer_flow(struct flow_cls_offload *f,
else
in_mdev = priv->mdev;
flow_flags |= BIT(MLX5E_TC_FLOW_FLAG_PEER);
parse_attr = flow->attr->parse_attr;
peer_flow = __mlx5e_add_fdb_flow(peer_priv, f, flow_flags,
parse_attr->filter_dev,

View File

@ -3986,7 +3986,7 @@ static void esw_offloads_steering_cleanup(struct mlx5_eswitch *esw)
mutex_destroy(&esw->fdb_table.offloads.vports.lock);
}
static void esw_vfs_changed_event_handler(struct mlx5_eswitch *esw)
static void esw_changed_event_handler(struct mlx5_eswitch *esw)
{
struct mlx5_esw_pf_info host_pf_info;
u16 new_num_vfs;
@ -3999,6 +3999,11 @@ static void esw_vfs_changed_event_handler(struct mlx5_eswitch *esw)
host_pf_info = mlx5_esw_get_host_pf_info(esw->dev, out);
new_num_vfs = host_pf_info.num_of_vfs;
if (host_pf_info.pf_disabled) {
mlx5_sf_table_esw_changed_event_handler(esw->dev);
mlx5_sf_hw_table_esw_changed_event_handler(esw->dev);
}
if (new_num_vfs == esw->esw_funcs.num_vfs || host_pf_info.pf_disabled)
goto free;
@ -4091,8 +4096,7 @@ int mlx5_esw_funcs_changed_handler(struct notifier_block *nb,
esw_funcs = mlx5_nb_cof(nb, struct mlx5_esw_functions, nb);
esw = container_of(esw_funcs, struct mlx5_eswitch, esw_funcs);
ret = mlx5_esw_add_work(esw, esw_vfs_changed_event_handler,
GFP_ATOMIC);
ret = mlx5_esw_add_work(esw, esw_changed_event_handler, GFP_ATOMIC);
if (ret)
return NOTIFY_DONE;

View File

@ -561,3 +561,32 @@ bool mlx5_sf_table_empty(const struct mlx5_core_dev *dev)
return xa_empty(&table->function_ids);
}
void mlx5_sf_table_esw_changed_event_handler(struct mlx5_core_dev *dev)
{
struct mlx5_sf_table *table = dev->priv.sf_table;
unsigned long index;
struct mlx5_sf *sf;
trace_mlx5_sf_host_pf_disabled(dev);
if (!table)
return;
mutex_lock(&table->sf_state_lock);
xa_for_each(&table->function_ids, index, sf) {
if (!sf->controller)
continue;
if (sf->hw_state == MLX5_VHCA_STATE_IN_USE)
sf->hw_state = MLX5_VHCA_STATE_ACTIVE;
else if (sf->hw_state == MLX5_VHCA_STATE_TEARDOWN_REQUEST)
sf->hw_state = MLX5_VHCA_STATE_ALLOCATED;
else
continue;
trace_mlx5_sf_update_state(table->dev, sf->port_index,
sf->controller, sf->hw_fn_id,
sf->hw_state);
}
mutex_unlock(&table->sf_state_lock);
}

View File

@ -11,6 +11,14 @@
#include <linux/mlx5/driver.h>
#include "sf/vhca_event.h"
TRACE_EVENT(mlx5_sf_host_pf_disabled,
TP_PROTO(const struct mlx5_core_dev *dev),
TP_ARGS(dev),
TP_STRUCT__entry(__string(devname, dev_name(dev->device))),
TP_fast_assign(__assign_str(devname);),
TP_printk("(%s)\n", __get_str(devname))
);
TRACE_EVENT(mlx5_sf_add,
TP_PROTO(const struct mlx5_core_dev *dev,
unsigned int port_index,

View File

@ -459,3 +459,25 @@ bool mlx5_sf_hw_table_supported(const struct mlx5_core_dev *dev)
{
return !!dev->priv.sf_hw_table;
}
void mlx5_sf_hw_table_esw_changed_event_handler(struct mlx5_core_dev *dev)
{
struct mlx5_sf_hw_table *table;
struct mlx5_sf_hwc_table *hwc;
int i;
table = dev->priv.sf_hw_table;
if (!table)
return;
mutex_lock(&table->table_lock);
hwc = &table->hwc[MLX5_SF_HWC_EXT_HOST];
for (i = 0; i < hwc->max_fn; i++) {
struct mlx5_sf_hw *sf_hw;
sf_hw = &hwc->sfs[i];
if (sf_hw->allocated && sf_hw->pending_delete)
mlx5_sf_hw_table_hwc_sf_free(dev, hwc, i);
}
mutex_unlock(&table->table_lock);
}

View File

@ -15,12 +15,14 @@ void mlx5_sf_hw_table_cleanup(struct mlx5_core_dev *dev);
int mlx5_sf_hw_notifier_init(struct mlx5_core_dev *dev);
void mlx5_sf_hw_notifier_cleanup(struct mlx5_core_dev *dev);
void mlx5_sf_hw_table_destroy(struct mlx5_core_dev *dev);
void mlx5_sf_hw_table_esw_changed_event_handler(struct mlx5_core_dev *dev);
int mlx5_sf_notifiers_init(struct mlx5_core_dev *dev);
int mlx5_sf_table_init(struct mlx5_core_dev *dev);
void mlx5_sf_notifiers_cleanup(struct mlx5_core_dev *dev);
void mlx5_sf_table_cleanup(struct mlx5_core_dev *dev);
bool mlx5_sf_table_empty(const struct mlx5_core_dev *dev);
void mlx5_sf_table_esw_changed_event_handler(struct mlx5_core_dev *dev);
int mlx5_devlink_sf_port_new(struct devlink *devlink,
const struct devlink_port_new_attrs *add_attr,
@ -60,6 +62,11 @@ static inline void mlx5_sf_hw_table_destroy(struct mlx5_core_dev *dev)
{
}
static inline void
mlx5_sf_hw_table_esw_changed_event_handler(struct mlx5_core_dev *dev)
{
}
static inline int mlx5_sf_notifiers_init(struct mlx5_core_dev *dev)
{
return 0;
@ -83,6 +90,11 @@ static inline bool mlx5_sf_table_empty(const struct mlx5_core_dev *dev)
return true;
}
static inline void
mlx5_sf_table_esw_changed_event_handler(struct mlx5_core_dev *dev)
{
}
#endif
#endif

View File

@ -2748,8 +2748,8 @@ static bool stmmac_xdp_xmit_zc(struct stmmac_priv *priv, u32 queue, u32 budget)
meta_req.set_ic = &set_ic;
meta_req.tbs = tx_q->tbs;
meta_req.edesc = &tx_q->dma_entx[entry];
xsk_tx_metadata_request(meta, &stmmac_xsk_tx_metadata_ops,
&meta_req);
xsk_tx_metadata_request(pool, &meta,
&stmmac_xsk_tx_metadata_ops, &meta_req);
if (set_ic) {
tx_q->tx_count_frames = 0;
stmmac_set_tx_ic(priv, tx_desc);
@ -4134,6 +4134,15 @@ static int __stmmac_open(struct net_device *dev,
dma_conf->tx_queue[i].tbs = priv->dma_conf.tx_queue[i].tbs;
memcpy(&priv->dma_conf, dma_conf, sizeof(*dma_conf));
/* The PHY is suspended when the interface is reopened without
* disconnecting the PHY, e.g. on MTU change. IEEE 802.3 allows PHYs
* to stop their receive clock while powered down, but the DMA
* software reset in stmmac_hw_setup() requires a running receive
* clock, and phylink_start() below resumes the PHY only after the
* hardware setup. Resume a suspended PHY here first.
*/
phylink_prepare_resume(priv->phylink);
stmmac_reset_queues_param(priv);
ret = stmmac_hw_setup(dev);

View File

@ -35,25 +35,11 @@ static void ovpn_priv_free(struct net_device *net)
static int ovpn_mp_alloc(struct ovpn_priv *ovpn)
{
struct in_device *dev_v4;
int i;
if (ovpn->mode != OVPN_MODE_MP)
return 0;
dev_v4 = __in_dev_get_rtnl(ovpn->dev);
if (dev_v4) {
/* disable redirects as Linux gets confused by ovpn
* handling same-LAN routing.
* This happens because a multipeer interface is used as
* relay point between hosts in the same subnet, while
* in a classic LAN this would not be needed because the
* two hosts would be able to talk directly.
*/
IN_DEV_CONF_SET(dev_v4, SEND_REDIRECTS, false);
IPV4_DEVCONF_ALL(dev_net(ovpn->dev), SEND_REDIRECTS) = false;
}
/* the peer container is fairly large, therefore we allocate it only in
* MP mode
*/
@ -97,9 +83,38 @@ static void ovpn_net_uninit(struct net_device *dev)
gro_cells_destroy(&ovpn->gro_cells);
}
static int ovpn_net_open(struct net_device *dev)
{
struct ovpn_priv *ovpn = netdev_priv(dev);
struct in_device *dev_v4;
/* the IPv4 in_device (and thus its config) is recreated whenever the
* interface is moved to a new netns, so redirects must be disabled on
* every bring-up rather than once at creation time, otherwise the
* setting is silently lost after such a move
*/
if (ovpn->mode == OVPN_MODE_MP) {
dev_v4 = __in_dev_get_rtnl(dev);
if (dev_v4) {
/* disable redirects as Linux gets confused by ovpn
* handling same-LAN routing.
* This happens because a multipeer interface is used as
* relay point between hosts in the same subnet, while
* in a classic LAN this would not be needed because the
* two hosts would be able to talk directly.
*/
IN_DEV_CONF_SET(dev_v4, SEND_REDIRECTS, false);
IPV4_DEVCONF_ALL(dev_net(dev), SEND_REDIRECTS) = false;
}
}
return 0;
}
static const struct net_device_ops ovpn_netdev_ops = {
.ndo_init = ovpn_net_init,
.ndo_uninit = ovpn_net_uninit,
.ndo_open = ovpn_net_open,
.ndo_start_xmit = ovpn_net_xmit,
};
@ -183,6 +198,7 @@ static int ovpn_newlink(struct net_device *dev,
struct ovpn_priv *ovpn = netdev_priv(dev);
struct nlattr **data = params->data;
enum ovpn_mode mode = OVPN_MODE_P2P;
int ret;
if (data && data[IFLA_OVPN_MODE]) {
mode = nla_get_u8(data[IFLA_OVPN_MODE]);
@ -207,7 +223,17 @@ static int ovpn_newlink(struct net_device *dev,
else
netif_carrier_off(dev);
return register_netdevice(dev);
ret = register_netdevice(dev);
if (ret < 0)
return ret;
return 0;
}
static size_t ovpn_get_size(const struct net_device *dev)
{
/* IFLA_OVPN_MODE */
return nla_total_size(sizeof(u8));
}
static int ovpn_fill_info(struct sk_buff *skb, const struct net_device *dev)
@ -228,13 +254,17 @@ static struct rtnl_link_ops ovpn_link_ops = {
.policy = ovpn_policy,
.maxtype = IFLA_OVPN_MAX,
.newlink = ovpn_newlink,
.get_size = ovpn_get_size,
.fill_info = ovpn_fill_info,
};
static int __init ovpn_init(void)
{
int err = rtnl_link_register(&ovpn_link_ops);
int err;
ovpn_tcp_init();
err = rtnl_link_register(&ovpn_link_ops);
if (err) {
pr_err("ovpn: can't register rtnl link ops: %d\n", err);
return err;
@ -246,8 +276,6 @@ static int __init ovpn_init(void)
goto unreg_rtnl;
}
ovpn_tcp_init();
return 0;
unreg_rtnl:

View File

@ -16,6 +16,14 @@ static const struct netlink_range_validation ovpn_a_peer_id_range = {
.max = 16777215ULL,
};
static const struct netlink_range_validation ovpn_a_peer_keepalive_interval_range = {
.max = 86400ULL,
};
static const struct netlink_range_validation ovpn_a_peer_keepalive_timeout_range = {
.max = 86400ULL,
};
static const struct netlink_range_validation ovpn_a_peer_tx_id_range = {
.max = 16777215ULL,
};
@ -68,8 +76,8 @@ const struct nla_policy ovpn_peer_nl_policy[OVPN_A_PEER_TX_ID + 1] = {
[OVPN_A_PEER_LOCAL_IPV4] = { .type = NLA_BE32, },
[OVPN_A_PEER_LOCAL_IPV6] = NLA_POLICY_EXACT_LEN(16),
[OVPN_A_PEER_LOCAL_PORT] = NLA_POLICY_MIN(NLA_BE16, 1),
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_interval_range),
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_timeout_range),
[OVPN_A_PEER_DEL_REASON] = NLA_POLICY_MAX(NLA_U32, 4),
[OVPN_A_PEER_VPN_RX_BYTES] = { .type = NLA_UINT, },
[OVPN_A_PEER_VPN_TX_BYTES] = { .type = NLA_UINT, },
@ -97,8 +105,8 @@ const struct nla_policy ovpn_peer_new_input_nl_policy[OVPN_A_PEER_TX_ID + 1] = {
[OVPN_A_PEER_VPN_IPV6] = NLA_POLICY_EXACT_LEN(16),
[OVPN_A_PEER_LOCAL_IPV4] = { .type = NLA_BE32, },
[OVPN_A_PEER_LOCAL_IPV6] = NLA_POLICY_EXACT_LEN(16),
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_interval_range),
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_timeout_range),
[OVPN_A_PEER_TX_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_tx_id_range),
};
@ -112,8 +120,8 @@ const struct nla_policy ovpn_peer_set_input_nl_policy[OVPN_A_PEER_TX_ID + 1] = {
[OVPN_A_PEER_VPN_IPV6] = NLA_POLICY_EXACT_LEN(16),
[OVPN_A_PEER_LOCAL_IPV4] = { .type = NLA_BE32, },
[OVPN_A_PEER_LOCAL_IPV6] = NLA_POLICY_EXACT_LEN(16),
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = { .type = NLA_U32, },
[OVPN_A_PEER_KEEPALIVE_INTERVAL] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_interval_range),
[OVPN_A_PEER_KEEPALIVE_TIMEOUT] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_keepalive_timeout_range),
[OVPN_A_PEER_TX_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &ovpn_a_peer_tx_id_range),
};

View File

@ -534,6 +534,12 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info)
*/
if (ret > 0)
ovpn_peer_hash_vpn_ip(peer);
/* if the remote endpoint was updated, the by_transp_addr hash bucket
* also needs to be refreshed, otherwise incoming packets from the new
* remote address would fail the lockless lookup
*/
if (attrs[OVPN_A_PEER_REMOTE_IPV4] || attrs[OVPN_A_PEER_REMOTE_IPV6])
ovpn_peer_hash_transp_addr(peer);
spin_unlock_bh(&ovpn->lock);
ovpn_peer_put(peer);

View File

@ -189,6 +189,9 @@ int ovpn_peer_reset_sockaddr(struct ovpn_peer *peer,
&(*__tbl1)[ovpn_get_hash_slot(*__tbl1, _key, _key_len)];\
})
static void __ovpn_peer_hash_transp_addr(struct ovpn_peer *peer,
const struct ovpn_bind *bind);
/**
* ovpn_peer_endpoints_update - update remote or local endpoint for peer
* @peer: peer to update the remote endpoint for
@ -196,7 +199,6 @@ int ovpn_peer_reset_sockaddr(struct ovpn_peer *peer,
*/
void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb)
{
struct hlist_nulls_head *nhead;
struct sockaddr_storage ss;
struct sockaddr_in6 *sa6;
bool reset_cache = false;
@ -220,9 +222,16 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb)
*/
local_ip = &ip_hdr(skb)->daddr;
sa = (struct sockaddr_in *)&ss;
sa->sin_family = AF_INET;
sa->sin_addr.s_addr = ip_hdr(skb)->saddr;
sa->sin_port = udp_hdr(skb)->source;
/* use a designated initializer so the sin_zero padding
* is zeroed (it ends up in the by_transp_addr hash key)
* without memset-ing the whole sockaddr_storage on the
* RX fast path
*/
*sa = (struct sockaddr_in) {
.sin_family = AF_INET,
.sin_addr.s_addr = ip_hdr(skb)->saddr,
.sin_port = udp_hdr(skb)->source,
};
salen = sizeof(*sa);
reset_cache = true;
break;
@ -248,11 +257,19 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb)
*/
local_ip = &ipv6_hdr(skb)->daddr;
sa6 = (struct sockaddr_in6 *)&ss;
sa6->sin6_family = AF_INET6;
sa6->sin6_addr = ipv6_hdr(skb)->saddr;
sa6->sin6_port = udp_hdr(skb)->source;
sa6->sin6_scope_id = ipv6_iface_scope_id(&ipv6_hdr(skb)->saddr,
skb->skb_iif);
/* use a designated initializer so the sin6_flowinfo
* padding is zeroed (it ends up in the by_transp_addr
* hash key) without memset-ing the whole
* sockaddr_storage on the RX fast path
*/
*sa6 = (struct sockaddr_in6) {
.sin6_family = AF_INET6,
.sin6_addr = ipv6_hdr(skb)->saddr,
.sin6_port = udp_hdr(skb)->source,
.sin6_scope_id =
ipv6_iface_scope_id(&ipv6_hdr(skb)->saddr,
skb->skb_iif),
};
salen = sizeof(*sa6);
reset_cache = true;
break;
@ -295,42 +312,25 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb)
ovpn_nl_peer_float_notify(peer, &ss);
/* rehashing is required only in MP mode as P2P has one peer
* only and thus there is no hashtable
* only and thus there is no hashtable.
*
* This function may be invoked concurrently, so re-read peer->bind
* under the proper locks and rehash against its current value.
*/
if (peer->ovpn->mode == OVPN_MODE_MP) {
spin_lock_bh(&peer->ovpn->lock);
spin_lock_bh(&peer->lock);
bind = rcu_dereference_protected(peer->bind,
lockdep_is_held(&peer->lock));
if (unlikely(!bind)) {
spin_unlock_bh(&peer->lock);
spin_unlock_bh(&peer->ovpn->lock);
return;
}
if (peer->ovpn->mode != OVPN_MODE_MP)
return;
/* This function may be invoked concurrently, therefore another
* float may have happened in parallel: perform rehashing
* using the peer->bind->remote directly as key
*/
switch (bind->remote.in4.sin_family) {
case AF_INET:
salen = sizeof(*sa);
break;
case AF_INET6:
salen = sizeof(*sa6);
break;
}
/* remove old hashing */
hlist_nulls_del_init_rcu(&peer->hash_entry_transp_addr);
/* re-add with new transport address */
nhead = ovpn_get_hash_head(peer->ovpn->peers->by_transp_addr,
&bind->remote, salen);
hlist_nulls_add_head_rcu(&peer->hash_entry_transp_addr, nhead);
spin_unlock_bh(&peer->lock);
spin_unlock_bh(&peer->ovpn->lock);
}
/* This function may be invoked concurrently, therefore another
* float may have happened in parallel: re-acquire the locks and
* rehash using the peer->bind->remote directly as key
*/
spin_lock_bh(&peer->ovpn->lock);
spin_lock_bh(&peer->lock);
bind = rcu_dereference_protected(peer->bind,
lockdep_is_held(&peer->lock));
__ovpn_peer_hash_transp_addr(peer, bind);
spin_unlock_bh(&peer->lock);
spin_unlock_bh(&peer->ovpn->lock);
return;
unlock:
spin_unlock_bh(&peer->lock);
@ -896,6 +896,83 @@ bool ovpn_peer_check_by_src(struct ovpn_priv *ovpn, struct sk_buff *skb,
return match;
}
/* Move @peer to the by_transp_addr bucket matching its current bind.
*
* Caller must hold both peer->ovpn->lock and peer->lock, and must have
* already dereferenced a valid (non-NULL) peer->bind, passed in as @bind.
*/
static void __ovpn_peer_hash_transp_addr(struct ovpn_peer *peer,
const struct ovpn_bind *bind)
{
struct sockaddr_storage sa = {};
struct hlist_nulls_head *nhead;
struct sockaddr_in6 *sa6;
struct sockaddr_in *sa4;
size_t salen;
lockdep_assert_held(&peer->ovpn->lock);
lockdep_assert_held(&peer->lock);
if (WARN_ON_ONCE(!bind))
return;
/* peer may have been concurrently removed between the caller's
* initial lookup and our acquisition of ovpn->lock; skip the
* rehash so we don't re-insert a removed peer
*/
if (unlikely(hlist_unhashed(&peer->hash_entry_id)))
return;
/* Build the hash key from the transport identity only
* (family/address/port), matching ovpn_peer_add_mp() and the lookup
* in ovpn_peer_get_by_transp_addr(). Hashing bind->remote directly
* would fold in sin6_scope_id (set on the float path but never by the
* lookup), scattering the peer into a bucket lookups cannot reach.
*/
switch (bind->remote.in4.sin_family) {
case AF_INET:
sa4 = (struct sockaddr_in *)&sa;
sa4->sin_family = AF_INET;
sa4->sin_addr.s_addr = bind->remote.in4.sin_addr.s_addr;
sa4->sin_port = bind->remote.in4.sin_port;
salen = sizeof(*sa4);
break;
case AF_INET6:
sa6 = (struct sockaddr_in6 *)&sa;
sa6->sin6_family = AF_INET6;
sa6->sin6_addr = bind->remote.in6.sin6_addr;
sa6->sin6_port = bind->remote.in6.sin6_port;
salen = sizeof(*sa6);
break;
default:
return;
}
/* remove old hashing (no-op if entry is not currently linked) */
hlist_nulls_del_init_rcu(&peer->hash_entry_transp_addr);
/* re-add with current transport address */
nhead = ovpn_get_hash_head(peer->ovpn->peers->by_transp_addr, &sa,
salen);
hlist_nulls_add_head_rcu(&peer->hash_entry_transp_addr, nhead);
}
void ovpn_peer_hash_transp_addr(struct ovpn_peer *peer)
{
struct ovpn_bind *bind;
lockdep_assert_held(&peer->ovpn->lock);
/* rehashing makes sense only in multipeer mode */
if (peer->ovpn->mode != OVPN_MODE_MP)
return;
spin_lock_bh(&peer->lock);
bind = rcu_dereference_protected(peer->bind,
lockdep_is_held(&peer->lock));
__ovpn_peer_hash_transp_addr(peer, bind);
spin_unlock_bh(&peer->lock);
}
void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer)
{
struct hlist_nulls_head *nhead;
@ -906,6 +983,13 @@ void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer)
if (peer->ovpn->mode != OVPN_MODE_MP)
return;
/* peer may have been concurrently removed between the caller's
* initial lookup and our acquisition of ovpn->lock; skip the
* rehash so we don't re-insert a removed peer
*/
if (hlist_unhashed(&peer->hash_entry_id))
return;
if (peer->vpn_addrs.ipv4.s_addr != htonl(INADDR_ANY)) {
/* remove potential old hashing */
hlist_nulls_del_init_rcu(&peer->hash_entry_addr4);
@ -1165,7 +1249,7 @@ static void ovpn_peer_release_p2p(struct ovpn_priv *ovpn, struct sock *sk,
}
if (sk) {
ovpn_sock = rcu_access_pointer(peer->sock);
ovpn_sock = rcu_dereference_bh(peer->sock);
if (!ovpn_sock || ovpn_sock->sk != sk) {
spin_unlock_bh(&ovpn->lock);
return;

View File

@ -150,6 +150,7 @@ struct ovpn_peer *ovpn_peer_get_by_id(struct ovpn_priv *ovpn, u32 peer_id);
struct ovpn_peer *ovpn_peer_get_by_dst(struct ovpn_priv *ovpn,
struct sk_buff *skb);
void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer);
void ovpn_peer_hash_transp_addr(struct ovpn_peer *peer);
bool ovpn_peer_check_by_src(struct ovpn_priv *ovpn, struct sk_buff *skb,
struct ovpn_peer *peer);

View File

@ -162,6 +162,15 @@ struct ovpn_socket *ovpn_socket_new(struct socket *sock, struct ovpn_peer *peer)
rcu_read_lock();
ovpn_sock = rcu_dereference_sk_user_data(sk);
if (ovpn_sock) {
/* something else filled the sk_user_data without
* setting the encap_type. Reject the socket.
*/
if (!type) {
ovpn_sock = ERR_PTR(-EBUSY);
rcu_read_unlock();
goto sock_release;
}
/* socket owned by another ovpn instance, we can't use it */
if (ovpn_sock->ovpn != peer->ovpn) {
ovpn_sock = ERR_PTR(-EBUSY);

View File

@ -53,15 +53,15 @@
#define MTK_GPHY_LED_RX_BLINK_SET (MTK_PHY_LED_BLINK_1000RX | \
MTK_PHY_LED_BLINK_100RX | \
MTK_PHY_LED_BLINK_10RX)
#define MTK_GPHY_LED_TX_BLINK_SET (MTK_PHY_LED_BLINK_1000RX | \
MTK_PHY_LED_BLINK_100RX | \
MTK_PHY_LED_BLINK_10RX)
#define MTK_GPHY_LED_TX_BLINK_SET (MTK_PHY_LED_BLINK_1000TX | \
MTK_PHY_LED_BLINK_100TX | \
MTK_PHY_LED_BLINK_10TX)
#define MTK_2P5GPHY_LED_ON_SET (MTK_PHY_LED_ON_LINK2500 | \
MTK_GPHY_LED_ON_SET)
#define MTK_2P5GPHY_LED_RX_BLINK_SET (MTK_PHY_LED_BLINK_2500RX | \
MTK_GPHY_LED_RX_BLINK_SET)
#define MTK_2P5GPHY_LED_TX_BLINK_SET (MTK_PHY_LED_BLINK_2500RX | \
#define MTK_2P5GPHY_LED_TX_BLINK_SET (MTK_PHY_LED_BLINK_2500TX | \
MTK_GPHY_LED_TX_BLINK_SET)
#define MTK_PHY_LED_STATE_FORCE_ON 0

View File

@ -1074,10 +1074,21 @@ static int tap_get_user_xdp(struct tap_queue *q, struct xdp_buff *xdp)
skb_reset_mac_header(skb);
skb->protocol = eth_hdr(skb)->h_proto;
rcu_read_lock();
tap = rcu_dereference(q->tap);
if (!tap) {
kfree_skb(skb);
rcu_read_unlock();
return 0;
}
skb->dev = tap->dev;
if (vnet_hdr_len) {
err = tun_vnet_hdr_to_skb(q->flags, skb, gso);
if (err)
if (err) {
rcu_read_unlock();
goto err_kfree;
}
}
/* Move network header to the right position for VLAN tagged packets */
@ -1085,15 +1096,8 @@ static int tap_get_user_xdp(struct tap_queue *q, struct xdp_buff *xdp)
vlan_get_protocol_and_depth(skb, skb->protocol, &depth) != 0)
skb_set_network_header(skb, depth);
rcu_read_lock();
tap = rcu_dereference(q->tap);
if (tap) {
skb->dev = tap->dev;
skb_probe_transport_header(skb);
dev_queue_xmit(skb);
} else {
kfree_skb(skb);
}
skb_probe_transport_header(skb);
dev_queue_xmit(skb);
rcu_read_unlock();
return 0;

View File

@ -386,11 +386,16 @@ static void tbnet_tear_down(struct tbnet *net, bool send_logout)
break;
}
tb_ring_stop(net->rx_ring.ring);
tb_ring_stop(net->tx_ring.ring);
tbnet_free_buffers(&net->rx_ring);
tbnet_free_buffers(&net->tx_ring);
/* Tear the paths down before stopping the rings. This mirrors
* tbnet_connected_work(), which enables the paths last so the
* Rx ring is primed before packets can arrive. Stopping a
* ring zeroes its descriptor base and tbnet_free_buffers()
* unmaps and frees the frame buffers, leaving anything still
* in flight with nowhere to drain to;
* __tb_path_deactivate_hop() then waits for the hop's
* 'pending' bit, which on some host routers never clears in
* that state.
*/
ret = tb_xdomain_disable_paths(net->xd,
net->local_transmit_path,
net->tx_ring.ring->hop,
@ -399,6 +404,11 @@ static void tbnet_tear_down(struct tbnet *net, bool send_logout)
if (ret)
netdev_warn(net->dev, "failed to disable DMA paths\n");
tb_ring_stop(net->rx_ring.ring);
tb_ring_stop(net->tx_ring.ring);
tbnet_free_buffers(&net->rx_ring);
tbnet_free_buffers(&net->tx_ring);
tb_xdomain_release_in_hopid(net->xd, net->remote_transmit_path);
net->remote_transmit_path = 0;
}
@ -925,12 +935,8 @@ static int tbnet_open(struct net_device *dev)
netif_carrier_off(dev);
flags = RING_FLAG_FRAME;
/* Only enable full E2E if the other end supports it too */
if (tbnet_e2e && net->svc->prtcstns & TBNET_E2E)
flags |= RING_FLAG_E2E;
ring = tb_ring_alloc_tx(xd->tb->nhi, -1, TBNET_RING_SIZE, flags);
ring = tb_ring_alloc_tx(xd->tb->nhi, -1, TBNET_RING_SIZE,
RING_FLAG_FRAME);
if (!ring) {
netdev_err(dev, "failed to allocate Tx ring\n");
return -ENOMEM;
@ -949,6 +955,11 @@ static int tbnet_open(struct net_device *dev)
sof_mask = BIT(TBIP_PDF_FRAME_START);
eof_mask = BIT(TBIP_PDF_FRAME_END);
flags = RING_FLAG_FRAME;
/* Only enable full E2E if the other end supports it too */
if (tbnet_e2e && net->svc->prtcstns & TBNET_E2E)
flags |= RING_FLAG_E2E;
ring = tb_ring_alloc_rx(xd->tb->nhi, -1, TBNET_RING_SIZE, flags,
net->tx_ring.ring->hop, sof_mask,
eof_mask, tbnet_start_poll, net);

View File

@ -1487,8 +1487,10 @@ ax88179_tx_fixup(struct usbnet *dev, struct sk_buff *skb, gfp_t flags)
headroom = skb_headroom(skb) - 8;
if ((dev->net->features & NETIF_F_SG) && skb_linearize(skb))
if ((dev->net->features & NETIF_F_SG) && skb_linearize(skb)) {
dev_kfree_skb_any(skb);
return NULL;
}
if ((skb_header_cloned(skb) || headroom < 0) &&
pskb_expand_head(skb, headroom < 0 ? 8 : 0, 0, GFP_ATOMIC)) {

View File

@ -490,6 +490,7 @@ static int ipheth_open(struct net_device *net)
if (retval)
return retval;
enable_delayed_work(&dev->carrier_work);
schedule_delayed_work(&dev->carrier_work, IPHETH_CARRIER_CHECK_TIMEOUT);
return retval;
}
@ -499,7 +500,11 @@ static int ipheth_close(struct net_device *net)
struct ipheth_device *dev = netdev_priv(net);
netif_stop_queue(net);
cancel_delayed_work_sync(&dev->carrier_work);
/* A TX URB can still complete with an error after this point and
* try to re-arm the carrier work. Disable it instead of cancelling
* it, so that such a schedule_delayed_work() is a no-op.
*/
disable_delayed_work_sync(&dev->carrier_work);
return 0;
}
@ -629,6 +634,10 @@ static int ipheth_probe(struct usb_interface *intf,
}
INIT_DELAYED_WORK(&dev->carrier_work, ipheth_carrier_check_work);
/* Armed only between ipheth_open() and ipheth_close(). Start out
* disabled so the enable/disable counts balance from the first open.
*/
disable_delayed_work(&dev->carrier_work);
retval = ipheth_alloc_urbs(dev);
if (retval) {

View File

@ -1794,7 +1794,7 @@ usbnet_probe(struct usb_interface *udev, const struct usb_device_id *prod)
*/
dev->hard_mtu = net->mtu + net->hard_header_len;
net->min_mtu = 0;
net->max_mtu = ETH_MAX_MTU;
net->max_mtu = net->mtu;
net->netdev_ops = &usbnet_netdev_ops;
net->watchdog_timeo = TX_TIMEOUT_JIFFIES;
@ -1804,6 +1804,7 @@ usbnet_probe(struct usb_interface *udev, const struct usb_device_id *prod)
// allow device-specific bind/init procedures
// NOTE net->name still not usable ...
if (info->bind) {
net->max_mtu = ETH_MAX_MTU;
status = info->bind(dev, udev);
if (status < 0)
goto out1;

View File

@ -2177,9 +2177,11 @@ ptp_ocp_devlink_info_get(struct devlink *devlink, struct devlink_info_req *req,
if (err)
return err;
snprintf(buf, sizeof(buf), "%.*s", OCP_BOARD_ID_LEN,
(const char *)bp->board_id);
err = devlink_info_version_fixed_put(req,
DEVLINK_INFO_VERSION_GENERIC_BOARD_ID,
bp->board_id);
buf);
if (err)
return err;

View File

@ -148,13 +148,16 @@ static int unregister_sba(struct ism_dev *ism)
if (ret && ret != ISM_ERROR)
return -EIO;
return 0;
}
static void ism_free_sba(struct ism_dev *ism)
{
dma_free_coherent(&ism->pdev->dev, PAGE_SIZE,
ism->sba, ism->sba_dma_addr);
ism->sba = NULL;
ism->sba_dma_addr = 0;
return 0;
}
static int unregister_ieq(struct ism_dev *ism)
@ -168,13 +171,16 @@ static int unregister_ieq(struct ism_dev *ism)
if (ret && ret != ISM_ERROR)
return -EIO;
return 0;
}
static void ism_free_ieq(struct ism_dev *ism)
{
dma_free_coherent(&ism->pdev->dev, PAGE_SIZE,
ism->ieq, ism->ieq_dma_addr);
ism->ieq = NULL;
ism->ieq_dma_addr = 0;
return 0;
}
static int ism_read_local_gid(struct dibs_dev *dibs)
@ -573,6 +579,7 @@ static int ism_dev_init(struct ism_dev *ism)
unreg_sba:
unregister_sba(ism);
ism_free_sba(ism);
free_irq:
free_irq(pci_irq_vector(pdev, 0), ism);
free_vectors:
@ -585,9 +592,13 @@ static void ism_dev_exit(struct ism_dev *ism)
{
struct pci_dev *pdev = ism->pdev;
/* ism will only generate new IRQs while ieq & sba are registered */
unregister_ieq(ism);
unregister_sba(ism);
/* drain ongoing irpt handlers */
free_irq(pci_irq_vector(pdev, 0), ism);
ism_free_ieq(ism);
ism_free_sba(ism);
pci_free_irq_vectors(pdev);
}

View File

@ -4710,6 +4710,9 @@ static int qeth_snmp_command(struct qeth_card *card, char __user *udata)
if (req_len > QETH_BUFSIZE)
return -EINVAL;
if (qinfo.udata_len < sizeof(struct qeth_snmp_ureq_hdr))
return -EINVAL;
iob = qeth_get_adapter_cmd(card, IPA_SETADP_SET_SNMP_CONTROL, req_len);
if (!iob)
return -ENOMEM;

View File

@ -1415,6 +1415,11 @@ static int qeth_l3_arp_query(struct qeth_card *card, char __user *udata)
rc = -EFAULT;
goto out;
}
if (qinfo.udata_len < QETH_QARP_ENTRIES_OFFSET) {
rc = -EINVAL;
goto out;
}
qinfo.udata = kzalloc(qinfo.udata_len, GFP_KERNEL);
if (!qinfo.udata) {
rc = -ENOMEM;

View File

@ -439,7 +439,7 @@ static inline void *dibs_get_priv(struct dibs_dev *dev,
/**
* dibs_dev_alloc() - allocate and reference device structure
*
* The following fields will be valid upon successful return: dev
* The following fields will be valid upon successful return: dev, lock
* NOTE: Use put_device(dibs_get_dev(@dibs)) to give up your reference instead
* of freeing @dibs @dev directly once you have successfully called this
* function.

View File

@ -300,9 +300,11 @@ struct hh_cache {
* We could use other alignment values, but we must maintain the
* relationship HH alignment <= LL alignment.
*/
#define LL_RESERVED_SPACE(dev) \
((((dev)->hard_header_len + READ_ONCE((dev)->needed_headroom)) \
#define LL_RESERVED_SPACE_EX(dev, hlen) \
((((hlen) + READ_ONCE((dev)->needed_headroom)) \
& ~(HH_DATA_MOD - 1)) + HH_DATA_MOD)
#define LL_RESERVED_SPACE(dev) \
LL_RESERVED_SPACE_EX(dev, (dev)->hard_header_len)
#define LL_RESERVED_SPACE_EXTRA(dev,extra) \
((((dev)->hard_header_len + READ_ONCE((dev)->needed_headroom) + (extra)) \
& ~(HH_DATA_MOD - 1)) + HH_DATA_MOD)
@ -3531,11 +3533,6 @@ static inline bool dev_validate_header(const struct net_device *dev,
if (len < dev->min_header_len)
return false;
if (capable(CAP_SYS_RAWIO)) {
memset(ll_header + len, 0, dev->hard_header_len - len);
return true;
}
if (dev->header_ops && dev->header_ops->validate)
return dev->header_ops->validate(ll_header, len);

View File

@ -244,8 +244,8 @@ extern void ip_set_type_unregister(struct ip_set_type *set_type);
/* A generic IP set */
struct ip_set {
/* For call_cru in destroy */
struct rcu_head rcu;
/* for set destruction */
struct rcu_work rwork;
/* The name of the set */
char name[IPSET_MAXNAMELEN];
/* Lock protecting the set data */
@ -273,7 +273,7 @@ struct ip_set {
/* Number of elements (vs timeout) */
u32 elements;
/* Size of the dynamic extensions (vs timeout) */
size_t ext_size;
atomic64_t ext_size;
/* Element data size */
size_t dsize;
/* Offsets to extensions in elements */

View File

@ -405,8 +405,8 @@ static inline struct inet6_dev *in6_dev_get(const struct net_device *dev)
rcu_read_lock();
idev = rcu_dereference(dev->ip6_ptr);
if (idev)
refcount_inc(&idev->refcnt);
if (idev && !refcount_inc_not_zero(&idev->refcnt))
idev = NULL;
rcu_read_unlock();
return idev;
}

View File

@ -25,9 +25,7 @@
#include <linux/netfilter.h> /* for union nf_inet_addr */
#include <linux/ip.h>
#include <linux/ipv6.h> /* for struct ipv6hdr */
#include <net/route.h>
#include <net/ipv6.h>
#include <net/ip6_fib.h>
#if IS_ENABLED(CONFIG_NF_CONNTRACK)
#include <net/netfilter/nf_conntrack.h>
#endif
@ -2062,7 +2060,7 @@ static inline bool ip_vs_conn_use_hash2(struct ip_vs_conn *cp)
void ip_vs_nat_icmp(struct sk_buff *skb, struct ip_vs_protocol *pp,
struct ip_vs_conn *cp, int dir, unsigned int toff,
bool has_ports);
bool has_ports, struct ip_vs_iphdr *ciph);
#ifdef CONFIG_IP_VS_IPV6
void ip_vs_nat_icmp_v6(struct sk_buff *skb, struct ip_vs_protocol *pp,
@ -2095,30 +2093,23 @@ static inline __wsum ip_vs_check_diff2(__be16 old, __be16 new, __wsum oldsum)
return csum_partial(diff, sizeof(diff), oldsum);
}
static inline bool ip_vs_checksum_needed(struct sk_buff *skb, int af)
static inline bool ip_vs_checksum_needed(struct sk_buff *skb)
{
/* Checksum unnecessary or already validated? */
if (skb_csum_unnecessary(skb))
return false;
/* LOCAL_OUT ? */
if (!skb->dev || skb->dev->flags & IFF_LOOPBACK)
/* Locally generated ? */
if (!skb->dev)
return false;
/* !LOCAL_IN (FORWARD) ? */
if (af == AF_INET6) {
if (!(dst_rt6_info(skb_dst(skb))->rt6i_flags & RTF_LOCAL))
return false;
} else {
if (!(skb_rtable(skb)->rt_flags & RTCF_LOCAL))
return false;
}
return true;
}
static inline bool ip_vs_checksum_common_check(struct sk_buff *skb,
int offset, int proto, int af)
{
if (!ip_vs_checksum_needed(skb, af))
if (!ip_vs_checksum_needed(skb))
return true;
/* Validate csum even for FORWARD */
return !nf_checksum(skb, NF_INET_LOCAL_IN, offset, proto, af);
}

View File

@ -205,7 +205,7 @@ __libeth_xsk_xmit_fill_buf_md(const struct xdp_desc *xdesc,
BUILD_BUG_ON(!__builtin_constant_p(tmo == libeth_xsktmo));
tmo = tmo == libeth_xsktmo ? &__libeth_xsktmo : tmo;
xsk_tx_metadata_request(ctx.meta, tmo, &desc);
xsk_tx_metadata_request(sq->pool, &ctx.meta, tmo, &desc);
return desc;
}

View File

@ -99,6 +99,7 @@ struct Qdisc {
struct hlist_node hash;
u32 handle;
u32 parent;
int depth;
struct netdev_queue *dev_queue;

View File

@ -141,45 +141,16 @@ INDIRECT_CALLABLE_DECLARE(void xsk_destruct_skb(struct sk_buff *));
static inline void xsk_tx_metadata_to_compl(struct xsk_tx_metadata *meta,
struct xsk_tx_metadata_compl *compl)
{
compl->tx_timestamp = NULL;
if (!meta)
return;
if (meta->flags & XDP_TXMD_FLAGS_TIMESTAMP)
compl->tx_timestamp = &meta->completion.tx_timestamp;
else
compl->tx_timestamp = NULL;
}
/* we can only arrive here if the completion timestamp has been
* requested via XDP_TXMD_FLAGS_TIMESTAMP, see xsk_tx_metadata_request
*/
/**
* xsk_tx_metadata_request - Evaluate AF_XDP TX metadata at submission
* and call appropriate xsk_tx_metadata_ops operation.
* @meta: pointer to AF_XDP metadata area
* @ops: pointer to struct xsk_tx_metadata_ops
* @priv: pointer to driver-private aread
*
* This function should be called by the networking device when
* it prepares AF_XDP egress packet.
*/
static inline void xsk_tx_metadata_request(const struct xsk_tx_metadata *meta,
const struct xsk_tx_metadata_ops *ops,
void *priv)
{
if (!meta)
return;
if (ops->tmo_request_launch_time)
if (meta->flags & XDP_TXMD_FLAGS_LAUNCH_TIME)
ops->tmo_request_launch_time(meta->request.launch_time,
priv);
if (ops->tmo_request_timestamp)
if (meta->flags & XDP_TXMD_FLAGS_TIMESTAMP)
ops->tmo_request_timestamp(priv);
if (ops->tmo_request_checksum)
if (meta->flags & XDP_TXMD_FLAGS_CHECKSUM)
ops->tmo_request_checksum(meta->request.csum_start,
meta->request.csum_offset, priv);
compl->tx_timestamp = &meta->completion.tx_timestamp;
}
/**
@ -231,12 +202,6 @@ static inline void xsk_tx_metadata_to_compl(struct xsk_tx_metadata *meta,
{
}
static inline void xsk_tx_metadata_request(struct xsk_tx_metadata *meta,
const struct xsk_tx_metadata_ops *ops,
void *priv)
{
}
static inline void xsk_tx_metadata_complete(struct xsk_tx_metadata_compl *compl,
const struct xsk_tx_metadata_ops *ops,
void *priv)

View File

@ -245,7 +245,7 @@ static inline void *xsk_buff_raw_get_data(struct xsk_buff_pool *pool, u64 addr)
* details.
*
* Return: new &xdp_desc_ctx struct containing desc's DMA address and metadata
* pointer, if it is present and valid (initialized to %NULL otherwise).
* pointer, if it is present (initialized to %NULL otherwise).
*/
static inline struct xdp_desc_ctx
xsk_buff_raw_get_ctx(const struct xsk_buff_pool *pool, u64 addr)
@ -260,24 +260,70 @@ xsk_buff_raw_get_ctx(const struct xsk_buff_pool *pool, u64 addr)
0)
static inline bool
xsk_buff_valid_tx_metadata(const struct xsk_tx_metadata *meta)
xsk_buff_valid_tx_metadata(const struct xsk_buff_pool *pool,
const struct xsk_tx_metadata *meta, u64 *flags)
{
return !(meta->flags & ~XDP_TXMD_FLAGS_VALID);
*flags = READ_ONCE(meta->flags);
if (*flags & XDP_TXMD_FLAGS_LAUNCH_TIME)
if (pool->tx_metadata_len <
offsetofend(struct xsk_tx_metadata, request.launch_time))
return false;
return !(*flags & ~XDP_TXMD_FLAGS_VALID);
}
/**
* xsk_tx_metadata_request - Evaluate AF_XDP TX metadata at submission
* and call appropriate xsk_tx_metadata_ops operation.
* @pool: pointer to AF_XDP buffer pool, used to validate the metadata
* @pmeta: pointer to pointer to AF_XDP metadata area
* @ops: pointer to struct xsk_tx_metadata_ops
* @priv: pointer to driver-private area
*
* This function should be called by the networking device when
* it prepares AF_XDP egress packet.
*/
static inline void
xsk_tx_metadata_request(const struct xsk_buff_pool *pool,
struct xsk_tx_metadata **pmeta,
const struct xsk_tx_metadata_ops *ops, void *priv)
{
const struct xsk_tx_metadata *meta = *pmeta;
u64 flags;
if (!meta)
return;
if (unlikely(!xsk_buff_valid_tx_metadata(pool, meta, &flags))) {
*pmeta = NULL;
return; /* no way to signal the error to the user */
}
if (ops->tmo_request_launch_time)
if (flags & XDP_TXMD_FLAGS_LAUNCH_TIME)
ops->tmo_request_launch_time(
READ_ONCE(meta->request.launch_time), priv);
if (ops->tmo_request_timestamp)
if (flags & XDP_TXMD_FLAGS_TIMESTAMP)
ops->tmo_request_timestamp(priv);
if (ops->tmo_request_checksum)
if (flags & XDP_TXMD_FLAGS_CHECKSUM)
ops->tmo_request_checksum(
READ_ONCE(meta->request.csum_start),
READ_ONCE(meta->request.csum_offset), priv);
if (!(flags & XDP_TXMD_FLAGS_TIMESTAMP))
*pmeta = NULL;
}
static inline struct xsk_tx_metadata *
__xsk_buff_get_metadata(const struct xsk_buff_pool *pool, void *data)
{
struct xsk_tx_metadata *meta;
if (!pool->tx_metadata_len)
return NULL;
meta = data - pool->tx_metadata_len;
if (unlikely(!xsk_buff_valid_tx_metadata(meta)))
return NULL; /* no way to signal the error to the user */
return meta;
return data - pool->tx_metadata_len;
}
static inline struct xsk_tx_metadata *
@ -469,11 +515,20 @@ xsk_buff_raw_get_ctx(const struct xsk_buff_pool *pool, u64 addr)
return (struct xdp_desc_ctx){ };
}
static inline bool xsk_buff_valid_tx_metadata(struct xsk_tx_metadata *meta)
static inline bool
xsk_buff_valid_tx_metadata(const struct xsk_buff_pool *pool,
const struct xsk_tx_metadata *meta, u64 *flags)
{
return false;
}
static inline void
xsk_tx_metadata_request(const struct xsk_buff_pool *pool,
struct xsk_tx_metadata **pmeta,
const struct xsk_tx_metadata_ops *ops, void *priv)
{
}
static inline struct xsk_tx_metadata *
__xsk_buff_get_metadata(const struct xsk_buff_pool *pool, void *data)
{

View File

@ -710,7 +710,7 @@ int vcc_setsockopt(struct socket *sock, int level, int optname,
sockptr_t optval, unsigned int optlen)
{
struct atm_vcc *vcc;
unsigned long value;
int value;
int error;
if (__SO_LEVEL_MATCH(optname, level) && optlen != __SO_SIZE(optname))
@ -722,8 +722,10 @@ int vcc_setsockopt(struct socket *sock, int level, int optname,
{
struct atm_qos qos;
if (copy_from_sockptr(&qos, optval, sizeof(qos)))
return -EFAULT;
error = copy_safe_from_sockptr(&qos, sizeof(qos), optval,
optlen);
if (error)
return error;
error = check_qos(&qos);
if (error)
return error;
@ -737,8 +739,10 @@ int vcc_setsockopt(struct socket *sock, int level, int optname,
return 0;
}
case SO_SETCLP:
if (copy_from_sockptr(&value, optval, sizeof(value)))
return -EFAULT;
error = copy_safe_from_sockptr(&value, sizeof(value), optval,
optlen);
if (error)
return error;
if (value)
vcc->atm_options |= ATM_ATMOPT_CLP;
else

View File

@ -224,11 +224,9 @@ static struct sk_buff *br_mrp_alloc_test_skb(struct br_mrp *mrp,
sub_opt = skb_put(skb, sizeof(*sub_opt));
memset(sub_opt, 0x0, sizeof(*sub_opt));
sub_tlv = skb_put(skb, sizeof(*sub_tlv));
sub_tlv->type = BR_MRP_SUB_TLV_HEADER_TEST_AUTO_MGR;
/* 32 bit alligment shall be ensured therefore add 2 bytes */
skb_put(skb, MRP_OPT_PADDING);
sub_tlv = skb_put_zero(skb, sizeof(*sub_tlv) + MRP_OPT_PADDING);
sub_tlv->type = BR_MRP_SUB_TLV_HEADER_TEST_AUTO_MGR;
}
br_mrp_skb_tlv(skb, BR_MRP_TLV_HEADER_END, 0x0);

View File

@ -41,11 +41,25 @@ ebt_nflog_tg(struct sk_buff *skb, const struct xt_action_param *par)
static int ebt_nflog_tg_check(const struct xt_tgchk_param *par)
{
struct ebt_nflog_info *info = par->targinfo;
int ret;
if (info->flags & ~EBT_NFLOG_MASK)
return -EINVAL;
info->prefix[EBT_NFLOG_PREFIX_SIZE - 1] = '\0';
return 0;
ret = nf_logger_find_get(par->family, NF_LOG_TYPE_ULOG);
if (ret != 0 && !par->nft_compat) {
request_module("%s", "nfnetlink_log");
ret = nf_logger_find_get(par->family, NF_LOG_TYPE_ULOG);
}
return ret;
}
static void ebt_nflog_tg_destroy(const struct xt_tgdtor_param *par)
{
nf_logger_put(par->family, NF_LOG_TYPE_ULOG);
}
static struct xt_target ebt_nflog_tg_reg __read_mostly = {
@ -54,6 +68,7 @@ static struct xt_target ebt_nflog_tg_reg __read_mostly = {
.family = NFPROTO_BRIDGE,
.target = ebt_nflog_tg,
.checkentry = ebt_nflog_tg_check,
.destroy = ebt_nflog_tg_destroy,
.targetsize = sizeof(struct ebt_nflog_info),
.me = THIS_MODULE,
};

View File

@ -712,6 +712,9 @@ zerocopy_fill_skb_from_devmem(struct sk_buff *skb, struct iov_iter *from,
size_t virt_addr, size, off;
struct net_iov *niov;
if (i && skb_frags_readable(skb))
return -EFAULT;
/* Devmem filling works by taking an IOVEC from the user where the
* iov_addrs are interpreted as an offset in bytes into the dma-buf to
* send from. We do not support other iter types.

View File

@ -11494,6 +11494,7 @@ int register_netdevice(struct net_device *dev)
* Prevent userspace races by waiting until the network
* device is fully setup before sending notifications.
*/
netdev_uevent_add(dev);
if (!(dev->rtnl_link_ops && dev->rtnl_link_initializing))
rtmsg_ifinfo(RTM_NEWLINK, dev, ~0U, GFP_KERNEL, 0, NULL);
@ -12435,6 +12436,7 @@ void unregister_netdevice_many_notify(struct list_head *head,
dev_tcx_uninstall(dev);
dev_xdp_uninstall(dev);
dev_memory_provider_uninstall(dev);
netdev_work_cancel_all(dev);
netdev_unlock_ops(dev);
bpf_dev_bound_netdev_unregister(dev);

View File

@ -179,6 +179,7 @@ enum netdev_work_core {
void __netdev_work_core_sched(struct net_device *dev, unsigned long event);
unsigned long
__netdev_work_core_cancel(struct net_device *dev, unsigned long mask);
void netdev_work_cancel_all(struct net_device *dev);
void __dev_notify_flags(struct net_device *dev, unsigned int old_flags,
unsigned int gchanges, u32 portid,

View File

@ -2334,6 +2334,9 @@ int netdev_register_kobject(struct net_device *ndev)
*groups++ = &wireless_group;
#endif /* CONFIG_SYSFS */
/* Hold back the KOBJ_ADD uevent until the device is listed. */
dev_set_uevent_suppress(dev, 1);
error = device_add(dev);
if (error)
return error;
@ -2349,6 +2352,17 @@ int netdev_register_kobject(struct net_device *ndev)
return error;
}
/* Announce a fully registered device to userspace. This pairs with the uevent
* suppression from netdev_register_kobject().
*/
void netdev_uevent_add(struct net_device *ndev)
{
struct device *dev = &ndev->dev;
dev_set_uevent_suppress(dev, 0);
kobject_uevent(&dev->kobj, KOBJ_ADD);
}
/* Change owner for sysfs entries when moving network devices across network
* namespaces owned by different user namespaces.
*/

View File

@ -4,6 +4,7 @@
int __init netdev_kobject_init(void);
int netdev_register_kobject(struct net_device *);
void netdev_uevent_add(struct net_device *dev);
void netdev_unregister_kobject(struct net_device *);
int net_rx_queue_update_kobjects(struct net_device *, int old_num, int new_num);
int netdev_queue_update_kobjects(struct net_device *net,

View File

@ -31,6 +31,10 @@ static void netdev_work_enqueue(struct net_device *dev, unsigned long events,
return;
spin_lock_bh(&netdev_work_lock);
if (!dev_isalive(dev)) {
spin_unlock_bh(&netdev_work_lock);
return;
}
if (list_empty(&dev->work_node)) {
list_add_tail(&dev->work_node, &netdev_work_list);
netdev_hold(dev, &dev->work_tracker, GFP_ATOMIC);
@ -61,6 +65,18 @@ netdev_work_dequeue(struct net_device *dev, unsigned long *pending,
return events;
}
void netdev_work_cancel_all(struct net_device *dev)
{
spin_lock_bh(&netdev_work_lock);
dev->work_pending = 0;
dev->work_core_pending = 0;
if (!list_empty(&dev->work_node)) {
list_del_init(&dev->work_node);
netdev_put(dev, &dev->work_tracker);
}
spin_unlock_bh(&netdev_work_lock);
}
void netdev_work_sched(struct net_device *dev, unsigned long events)
{
netdev_work_enqueue(dev, events, 0);

View File

@ -779,7 +779,6 @@ bool sk_mc_loop(const struct sock *sk)
return inet6_test_bit(MC6_LOOP, sk);
#endif
}
WARN_ON_ONCE(1);
return true;
}
EXPORT_SYMBOL(sk_mc_loop);

View File

@ -871,7 +871,7 @@ struct xdp_frame *xdpf_clone(struct xdp_frame *xdpf)
headroom = xdpf->headroom + sizeof(*xdpf);
totalsize = headroom + xdpf->len;
if (unlikely(totalsize > PAGE_SIZE))
if (unlikely(totalsize > SKB_WITH_OVERHEAD(PAGE_SIZE)))
return NULL;
page = dev_alloc_page();
if (!page)

View File

@ -578,6 +578,7 @@ int devlink_nl_reload_doit(struct sk_buff *skb, struct genl_info *info)
action != DEVLINK_RELOAD_ACTION_DRIVER_REINIT) {
NL_SET_ERR_MSG_MOD(info->extack,
"Changing namespace is only supported for reinit action");
put_net(dest_net);
return -EOPNOTSUPP;
}
}

View File

@ -490,6 +490,34 @@ int ip_fib_check_default(__be32 gw, struct net_device *dev)
return -1;
}
static size_t fib_nexthop_nlmsg_size(const struct fib_nh_common *nhc,
bool skip_oif)
{
size_t nhsize = 0;
switch (nhc->nhc_gw_family) {
case AF_INET:
nhsize += nla_total_size(4); /* RTA_GATEWAY */
break;
case AF_INET6:
nhsize += nla_total_size(sizeof(struct rtvia) +
sizeof(struct in6_addr));
break;
}
if (!skip_oif && nhc->nhc_dev)
nhsize += nla_total_size(4); /* RTA_OIF */
if (nhc->nhc_lwtstate) {
/* RTA_ENCAP */
nhsize += lwtunnel_get_encap_size(nhc->nhc_lwtstate);
/* RTA_ENCAP_TYPE */
nhsize += nla_total_size(2);
}
return nhsize;
}
size_t fib_nlmsg_size(struct fib_info *fi)
{
size_t payload = NLMSG_ALIGN(sizeof(struct rtmsg))
@ -507,32 +535,35 @@ size_t fib_nlmsg_size(struct fib_info *fi)
payload += nla_total_size(4); /* RTA_NH_ID */
if (nhs) {
size_t nh_encapsize = 0;
/* Also handles the special case nhs == 1 */
/* each nexthop is packed in an attribute */
size_t nhsize = nla_total_size(sizeof(struct rtnexthop));
size_t mpsize = 0;
unsigned int i;
/* may contain flow and gateway attribute */
nhsize += 2 * nla_total_size(4);
/* grab encap info */
for (i = 0; i < fib_info_num_path(fi); i++) {
struct fib_nh_common *nhc = fib_info_nhc(fi, i);
size_t nhsize;
if (nhc->nhc_lwtstate) {
/* RTA_ENCAP_TYPE */
nh_encapsize += lwtunnel_get_encap_size(
nhc->nhc_lwtstate);
/* RTA_ENCAP */
nh_encapsize += nla_total_size(2);
nhsize = fib_nexthop_nlmsg_size(nhc, nhs != 1);
if (nhs != 1)
nhsize += NLA_ALIGN(sizeof(struct rtnexthop));
#ifdef CONFIG_IP_ROUTE_CLASSID
if (nhc->nhc_family == AF_INET) {
struct fib_nh *nh;
nh = container_of(nhc, struct fib_nh, nh_common);
if (nh->nh_tclassid)
nhsize += nla_total_size(4);
}
#endif
if (nhs == 1)
payload += nhsize;
else
mpsize += nhsize;
}
/* all nexthops are packed in a nested attribute */
payload += nla_total_size((nhs * nhsize) + nh_encapsize);
if (nhs != 1)
payload += nla_total_size(mpsize);
}
return payload;

View File

@ -943,11 +943,23 @@ static struct request_sock *inet_reqsk_clone(struct request_sock *req,
nreq->rsk_listener = sk;
/* We need not acquire fastopenq->lock
* because the child socket is locked in inet_csk_listen_stop().
*/
if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(nreq)->tfo_listener)
if (sk->sk_protocol == IPPROTO_TCP && tcp_rsk(nreq)->tfo_listener) {
struct fastopen_queue *fastopenq;
/* reqsk_fastopen_remove() will uncharge nreq->rsk_listener,
* that is @sk, so charge it here. Unlike the listener
* being closed, @sk is live and needs its lock.
*/
fastopenq = &inet_csk(sk)->icsk_accept_queue.fastopenq;
spin_lock_bh(&fastopenq->lock);
fastopenq->qlen++;
spin_unlock_bh(&fastopenq->lock);
/* We need not acquire fastopenq->lock
* because the child socket is locked in inet_csk_listen_stop().
*/
rcu_assign_pointer(tcp_sk(nreq->sk)->fastopen_rsk, nreq);
}
return nreq;
}

View File

@ -393,8 +393,8 @@ static struct inet_frag_queue *inet_frag_create(struct fqdir *fqdir,
*prev = ERR_PTR(-ENOMEM);
return NULL;
}
mod_timer(&q->timer, jiffies + fqdir->timeout);
spin_lock_bh(&q->lock);
*prev = rhashtable_lookup_get_insert_key(&fqdir->rhashtable, &q->key,
&q->node, f->rhash_params);
if (*prev) {
@ -402,13 +402,13 @@ static struct inet_frag_queue *inet_frag_create(struct fqdir *fqdir,
* we need to cancel what inet_frag_alloc()
* anticipated.
*/
int refs = 1;
q->flags |= INET_FRAG_COMPLETE;
inet_frag_kill(q, &refs);
inet_frag_putn(q, refs);
spin_unlock_bh(&q->lock);
inet_frag_putn(q, 2);
return NULL;
}
mod_timer(&q->timer, jiffies + fqdir->timeout);
spin_unlock_bh(&q->lock);
return q;
}

View File

@ -252,7 +252,7 @@ static void tcp_measure_rcv_mss(struct sock *sk, const struct sk_buff *skb)
struct tcp_sock *tp = tcp_sk(sk);
val = tcp_win_from_space(sk, sk->sk_rcvbuf);
tcp_set_window_clamp(sk, val);
WRITE_ONCE(tp->window_clamp, val);
if (tp->window_clamp < tp->rcvq_space.space)
tp->rcvq_space.space = tp->window_clamp;

View File

@ -178,17 +178,19 @@ static struct sk_buff *__skb_udp_tunnel_segment(struct sk_buff *skb,
int tnl_hlen = skb_inner_mac_header(skb) - skb_transport_header(skb);
bool remcsum, need_csum, offload_csum, gso_partial;
struct sk_buff *segs = ERR_PTR(-EINVAL);
struct udphdr *uh = udp_hdr(skb);
u16 mac_offset = skb->mac_header;
__be16 protocol = skb->protocol;
u16 mac_len = skb->mac_len;
int udp_offset, outer_hlen;
struct udphdr *uh;
__wsum partial;
bool need_ipsec;
if (unlikely(!pskb_may_pull(skb, tnl_hlen)))
goto out;
uh = udp_hdr(skb);
/* Adjust partial header checksum to negate old length.
* We cannot rely on the value contained in uh->len as it is
* possible that the actual value exceeds the boundaries of the

View File

@ -684,6 +684,9 @@ ip6ip6_err(struct sk_buff *skb, struct inet6_skb_parm *opt,
if (!skb2)
return 0;
/* Remove debris left by outer IPv6 stack. */
memset(IP6CB(skb2), 0, sizeof(*IP6CB(skb2)));
skb_dst_drop(skb2);
skb_pull(skb2, offset);
skb_reset_network_header(skb2);

View File

@ -988,13 +988,13 @@ int rt6_route_rcv(struct net_device *dev, u8 *opt, int len,
} else if (rinfo->prefix_len > 128) {
return -EINVAL;
} else if (rinfo->prefix_len > 64) {
if (rinfo->length < 2) {
/* RFC 4191: Length MUST be 3 when Prefix Length > 64 */
if (rinfo->length < 3)
return -EINVAL;
}
} else if (rinfo->prefix_len > 0) {
if (rinfo->length < 1) {
/* RFC 4191: Length MUST be 2 or 3 when Prefix Length > 0 */
if (rinfo->length < 2)
return -EINVAL;
}
}
pref = rinfo->route_pref;

View File

@ -415,6 +415,7 @@ void mac802154_beacon_worker(struct work_struct *work)
container_of(work, struct ieee802154_local, beacon_work.work);
struct cfg802154_beacon_request *beacon_req;
struct ieee802154_sub_if_data *sdata;
netdevice_tracker dev_tracker;
struct wpan_dev *wpan_dev;
u8 interval;
int ret;
@ -427,12 +428,14 @@ void mac802154_beacon_worker(struct work_struct *work)
}
sdata = IEEE802154_WPAN_DEV_TO_SUB_IF(beacon_req->wpan_dev);
netdev_hold(sdata->dev, &dev_tracker, GFP_ATOMIC);
/* Wait an arbitrary amount of time in case we cannot use the device */
if (local->suspended || !ieee802154_sdata_running(sdata)) {
rcu_read_unlock();
queue_delayed_work(local->mac_wq, &local->beacon_work,
msecs_to_jiffies(1000));
netdev_put(sdata->dev, &dev_tracker);
return;
}
@ -450,6 +453,7 @@ void mac802154_beacon_worker(struct work_struct *work)
if (interval < IEEE802154_ACTIVE_SCAN_DURATION)
queue_delayed_work(local->mac_wq, &local->beacon_work,
local->beacon_interval);
netdev_put(sdata->dev, &dev_tracker);
}
int mac802154_stop_beacons_locked(struct ieee802154_local *local,

View File

@ -24,12 +24,13 @@ void mptcp_fastopen_subflow_synack_set_params(struct mptcp_subflow_context *subf
sk = subflow->conn;
tp = tcp_sk(ssk);
subflow->is_mptfo = 1;
/* A valid TFO cookie does not guarantee SYN data. */
skb = skb_peek(&ssk->sk_receive_queue);
if (WARN_ON_ONCE(!skb))
if (!skb)
return;
subflow->is_mptfo = 1;
/* dequeue the skb from sk receive queue */
__skb_unlink(skb, &ssk->sk_receive_queue);
skb_ext_reset(skb);

View File

@ -50,6 +50,14 @@ static void mptcp_parse_option(const struct sk_buff *skb,
}
}
/* Only the MPC + ACK can be used with a RM_ADDR */
if (subopt == OPTION_MPTCP_MPC_ACK) {
if ((mp_opt->suboptions & ~OPTION_MPTCP_RM_ADDR) != 0)
break;
} else if (mp_opt->suboptions != 0) {
break;
}
/* Cfr RFC 8684 Section 3.3.0:
* If a checksum is present but its use had
* not been negotiated in the MP_CAPABLE handshake, the receiver MUST
@ -122,6 +130,11 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_MP_JOIN:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTION_MPTCP_RM_ADDR |
OPTION_MPTCP_PRIO)) != 0)
break;
if (opsize == TCPOLEN_MPTCP_MPJ_SYN) {
mp_opt->suboptions |= OPTION_MPTCP_MPJ_SYN;
mp_opt->backup = *ptr++ & MPTCPOPT_BACKUP;
@ -153,6 +166,14 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_DSS:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTION_MPTCP_ADD_ADDR |
OPTION_MPTCP_RM_ADDR |
OPTION_MPTCP_PRIO |
OPTION_MPTCP_FASTCLOSE |
OPTION_MPTCP_FAIL)) != 0)
break;
pr_debug("DSS\n");
ptr++;
@ -188,8 +209,14 @@ static void mptcp_parse_option(const struct sk_buff *skb,
* RFC 8684 Section 3.3.0 checks later in subflow_data_ready
*/
if (opsize != expected_opsize &&
opsize != expected_opsize + TCPOLEN_MPTCP_DSS_CHECKSUM)
opsize != expected_opsize + TCPOLEN_MPTCP_DSS_CHECKSUM) {
mp_opt->dsn64 = 0;
mp_opt->use_map = 0;
mp_opt->ack64 = 0;
mp_opt->use_ack = 0;
mp_opt->data_fin = 0;
break;
}
mp_opt->suboptions |= OPTION_MPTCP_DSS;
if (mp_opt->use_ack) {
@ -234,6 +261,12 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_ADD_ADDR:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTIONS_MPTCP_DSS |
OPTION_MPTCP_RM_ADDR |
OPTION_MPTCP_PRIO)) != 0)
break;
mp_opt->echo = (*ptr++) & MPTCP_ADDR_ECHO;
if (!mp_opt->echo) {
if (opsize == TCPOLEN_MPTCP_ADD_ADDR ||
@ -293,6 +326,14 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_RM_ADDR:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTION_MPTCP_MPC_ACK |
OPTIONS_MPTCP_MPJ |
OPTIONS_MPTCP_DSS |
OPTION_MPTCP_ADD_ADDR |
OPTION_MPTCP_PRIO)) != 0)
break;
if (opsize < TCPOLEN_MPTCP_RM_ADDR_BASE + 1 ||
opsize > TCPOLEN_MPTCP_RM_ADDR_BASE + MPTCP_RM_IDS_MAX)
break;
@ -307,6 +348,13 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_MP_PRIO:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTIONS_MPTCP_MPJ |
OPTIONS_MPTCP_DSS |
OPTION_MPTCP_ADD_ADDR |
OPTION_MPTCP_RM_ADDR)) != 0)
break;
if (opsize != TCPOLEN_MPTCP_PRIO)
break;
@ -316,6 +364,11 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_MP_FASTCLOSE:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTIONS_MPTCP_DSS |
OPTION_MPTCP_RST)) != 0)
break;
if (opsize != TCPOLEN_MPTCP_FASTCLOSE)
break;
@ -327,6 +380,11 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_RST:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTION_MPTCP_FAIL |
OPTION_MPTCP_FASTCLOSE)) != 0)
break;
if (opsize != TCPOLEN_MPTCP_RST)
break;
@ -342,6 +400,11 @@ static void mptcp_parse_option(const struct sk_buff *skb,
break;
case MPTCPOPT_MP_FAIL:
/* Can be used with a restricted number of other options */
if ((mp_opt->suboptions & ~(OPTIONS_MPTCP_DSS |
OPTION_MPTCP_RST)) != 0)
break;
if (opsize != TCPOLEN_MPTCP_FAIL)
break;
@ -1400,7 +1463,7 @@ void mptcp_write_options(struct tcphdr *th, __be32 *ptr, struct tcp_sock *tp,
* RM | C | C | C | P |------|------|------|------|
* PRIO | X | C | C | C | C |------|------|------|
* FAIL | X | X | C | X | X | X |------|------|
* FC | X | X | X | X | X | X | X |------|
* FC | X | X | P | X | X | X | X |------|
* RST | X | X | X | X | X | X | O | O |
* ------|------|------|------|------|------|------|------|------|
*

View File

@ -380,6 +380,7 @@ static void mptcp_pm_add_addr_timer(struct timer_list *timer)
struct mptcp_sock *msk = entry->sock;
struct sock *sk = (struct sock *)msk;
unsigned int timeout = 0;
bool retransmit;
pr_debug("msk=%p\n", msk);
@ -412,14 +413,15 @@ static void mptcp_pm_add_addr_timer(struct timer_list *timer)
entry->retrans_times++;
}
if (entry->retrans_times < ADD_ADDR_RETRANS_MAX)
retransmit = entry->retrans_times < ADD_ADDR_RETRANS_MAX;
if (retransmit)
timeout <<= entry->retrans_times;
else
timeout = 0;
spin_unlock_bh(&msk->pm.lock);
if (entry->retrans_times == ADD_ADDR_RETRANS_MAX)
if (!retransmit)
mptcp_pm_subflow_established(msk);
out:
@ -441,6 +443,9 @@ bool mptcp_pm_announced_alloc(struct mptcp_sock *msk,
lockdep_assert_held(&msk->pm.lock);
if (msk->pm.status & BIT(MPTCP_PM_DESTROYING))
return false;
add_entry = mptcp_pm_announced_lookup(msk, addr);
if (add_entry) {
if (WARN_ON_ONCE(mptcp_pm_is_kernel(msk)))
@ -1143,10 +1148,16 @@ void mptcp_pm_worker(struct mptcp_sock *msk)
void mptcp_pm_destroy(struct mptcp_sock *msk)
{
spin_lock_bh(&msk->pm.lock);
msk->pm.status |= BIT(MPTCP_PM_DESTROYING);
spin_unlock_bh(&msk->pm.lock);
mptcp_pm_free_announced_list(msk);
if (mptcp_pm_is_userspace(msk))
mptcp_userspace_pm_free_local_addr_list(msk);
/* Free the userspace local address list unconditionally: the socket
* can be reused (mptcp_disconnect()) and re-selected to a different PM
*/
mptcp_userspace_pm_free_local_addr_list(msk);
}
void mptcp_pm_data_reset(struct mptcp_sock *msk)

View File

@ -54,6 +54,10 @@ static int mptcp_userspace_pm_append_new_local_addr(struct mptcp_sock *msk,
bitmap_zero(id_bitmap, MPTCP_PM_MAX_ADDR_ID + 1);
spin_lock_bh(&msk->pm.lock);
if (msk->pm.status & BIT(MPTCP_PM_DESTROYING)) {
ret = -EINVAL;
goto append_err;
}
mptcp_for_each_userspace_pm_addr(msk, e) {
addr_match = mptcp_addresses_equal(&e->addr, &entry->addr, true);
if (addr_match && entry->addr.id == 0 && needs_id)

View File

@ -149,6 +149,12 @@ struct sock *__mptcp_nmpc_sk(struct mptcp_sock *msk)
static void mptcp_drop(struct sock *sk, struct sk_buff *skb)
{
/* The skb forward memory was already transferred to sk by
* mptcp_borrow_fwdmem(), even before setting the destructor.
*/
if (!skb->destructor)
sk_mem_reclaim(sk);
sk_drops_skbadd(sk, skb);
__kfree_skb(skb);
}

View File

@ -37,6 +37,7 @@
OPTION_MPTCP_MPC_ACK)
#define OPTIONS_MPTCP_MPJ (OPTION_MPTCP_MPJ_SYN | OPTION_MPTCP_MPJ_SYNACK | \
OPTION_MPTCP_MPJ_ACK)
#define OPTIONS_MPTCP_DSS (OPTION_MPTCP_DSS | OPTION_MPTCP_CSUMREQD)
/* MPTCP option subtypes */
#define MPTCPOPT_MP_CAPABLE 0
@ -189,9 +190,10 @@ enum mptcp_pm_status {
MPTCP_PM_ESTABLISHED,
MPTCP_PM_SUBFLOW_ESTABLISHED,
MPTCP_PM_ALREADY_ESTABLISHED, /* persistent status, set after ESTABLISHED event */
MPTCP_PM_MPC_ENDPOINT_ACCOUNTED /* persistent status, set after MPC local address is
* accounted int id_avail_bitmap
*/
MPTCP_PM_MPC_ENDPOINT_ACCOUNTED, /* persistent status, set after MPC local address is
* accounted int id_avail_bitmap
*/
MPTCP_PM_DESTROYING, /* To fence out PM list allocs */
};
enum mptcp_pm_type {

View File

@ -174,8 +174,6 @@ static int subflow_check_req(struct request_sock *req,
if (unlikely(listener->pm_listener))
return subflow_reset_req_endp(req, skb);
if (opt_mp_join)
return 0;
} else if (opt_mp_join) {
SUBFLOW_REQ_INC_STATS(req, MPTCP_MIB_JOINSYNRX);
@ -277,9 +275,6 @@ int mptcp_subflow_init_cookie_req(struct request_sock *req,
opt_mp_capable = !!(mp_opt.suboptions & OPTION_MPTCP_MPC_ACK);
opt_mp_join = !!(mp_opt.suboptions & OPTION_MPTCP_MPJ_ACK);
if (opt_mp_capable && opt_mp_join)
return -EINVAL;
if (opt_mp_capable && listener->request_mptcp) {
if (mp_opt.sndr_key == 0)
return -EINVAL;

View File

@ -461,6 +461,10 @@ static int ncsi_send_cmd_nl(struct sk_buff *msg, struct genl_info *info)
nca.req_flags = NCSI_REQ_FLAG_NETLINK_DRIVEN;
nca.info = info;
nca.payload = ntohs(hdr->length);
if (nca.payload > len - sizeof(*hdr)) {
ret = -EINVAL;
goto out_netlink;
}
nca.data = data + sizeof(*hdr);
ret = ncsi_xmit_cmd(&nca);

View File

@ -77,7 +77,7 @@ mtype_flush(struct ip_set *set)
mtype_ext_cleanup(set);
bitmap_zero(map->members, map->elements);
set->elements = 0;
set->ext_size = 0;
atomic64_set(&set->ext_size, 0);
}
/* Calculate the actual memory size of the set data */
@ -93,7 +93,7 @@ mtype_head(struct ip_set *set, struct sk_buff *skb)
{
const struct mtype *map = set->data;
struct nlattr *nested;
size_t memsize = mtype_memsize(map, set->dsize) + set->ext_size;
size_t memsize = mtype_memsize(map, set->dsize) + atomic64_read(&set->ext_size);
nested = nla_nest_start(skb, IPSET_ATTR_DATA);
if (!nested)

View File

@ -25,6 +25,7 @@
static LIST_HEAD(ip_set_type_list); /* all registered set types */
static DEFINE_MUTEX(ip_set_type_mutex); /* protects ip_set_type_list */
static DEFINE_RWLOCK(ip_set_ref_lock); /* protects the set refs */
static struct workqueue_struct *ipset_destroy_wq;
struct ip_set_net {
struct ip_set * __rcu *ip_set_list; /* all individual sets */
@ -350,7 +351,7 @@ ip_set_init_comment(struct ip_set *set, struct ip_set_comment *comment,
size_t len = ext->comment ? strlen(ext->comment) : 0;
if (unlikely(c)) {
set->ext_size -= sizeof(*c) + strlen(c->str) + 1;
atomic64_sub(sizeof(*c) + strlen(c->str) + 1, &set->ext_size);
rcu_assign_pointer(comment->c, NULL);
kfree_rcu(c, rcu);
}
@ -362,7 +363,7 @@ ip_set_init_comment(struct ip_set *set, struct ip_set_comment *comment,
if (unlikely(!c))
return;
strscpy(c->str, ext->comment, len + 1);
set->ext_size += sizeof(*c) + strlen(c->str) + 1;
atomic64_add(sizeof(*c) + strlen(c->str) + 1, &set->ext_size);
rcu_assign_pointer(comment->c, c);
}
EXPORT_SYMBOL_GPL(ip_set_init_comment);
@ -392,7 +393,7 @@ ip_set_comment_free(struct ip_set *set, void *ptr)
c = rcu_dereference_protected(comment->c, 1);
if (unlikely(!c))
return;
set->ext_size -= sizeof(*c) + strlen(c->str) + 1;
atomic64_sub(sizeof(*c) + strlen(c->str) + 1, &set->ext_size);
rcu_assign_pointer(comment->c, NULL);
kfree_rcu(c, rcu);
}
@ -1178,22 +1179,26 @@ ip_set_setname_policy[IPSET_ATTR_CMD_MAX + 1] = {
.len = IPSET_MAXNAMELEN - 1 },
};
/* In order to return quickly when destroying a single set, it is split
* into two stages:
* - Cancel garbage collector
* - Destroy the set itself via call_rcu()
*/
static void
ip_set_destroy_set_rcu(struct rcu_head *head)
destroy_and_free_set(struct ip_set *set)
{
struct ip_set *set = container_of(head, struct ip_set, rcu);
set->variant->destroy(set);
module_put(set->type->me);
kfree(set);
}
/* In order to return quickly when destroying a single set,
* destruction is done asynchronously via work queues.
*/
static void
ip_set_destroy_set_work(struct work_struct *work)
{
struct ip_set *set = container_of(to_rcu_work(work),
struct ip_set, rwork);
destroy_and_free_set(set);
}
static void
_destroy_all_sets(struct ip_set_net *inst)
{
@ -1283,7 +1288,8 @@ static int ip_set_destroy(struct sk_buff *skb, const struct nfnl_info *info,
/* Must wait for flush to be really finished */
rcu_barrier();
}
call_rcu(&s->rcu, ip_set_destroy_set_rcu);
INIT_RCU_WORK(&s->rwork, ip_set_destroy_set_work);
queue_rcu_work(ipset_destroy_wq, &s->rwork);
}
return 0;
out:
@ -2421,18 +2427,23 @@ static struct pernet_operations ip_set_net_ops = {
static int __init
ip_set_init(void)
{
int ret = register_pernet_subsys(&ip_set_net_ops);
int ret;
ipset_destroy_wq = alloc_ordered_workqueue("ipset_destroy_wq", 0);
if (!ipset_destroy_wq)
return -ENOMEM;
ret = register_pernet_subsys(&ip_set_net_ops);
if (ret) {
pr_err("ip_set: cannot register pernet_subsys.\n");
return ret;
goto out_wq;
}
ret = nfnetlink_subsys_register(&ip_set_netlink_subsys);
if (ret != 0) {
pr_err("ip_set: cannot register with nfnetlink.\n");
unregister_pernet_subsys(&ip_set_net_ops);
return ret;
goto out_wq;
}
ret = nf_register_sockopt(&so_set);
@ -2440,10 +2451,13 @@ ip_set_init(void)
pr_err("SO_SET registry failed: %d\n", ret);
nfnetlink_subsys_unregister(&ip_set_netlink_subsys);
unregister_pernet_subsys(&ip_set_net_ops);
return ret;
goto out_wq;
}
return 0;
out_wq:
destroy_workqueue(ipset_destroy_wq);
return ret;
}
static void __exit
@ -2453,9 +2467,7 @@ ip_set_fini(void)
nfnetlink_subsys_unregister(&ip_set_netlink_subsys);
unregister_pernet_subsys(&ip_set_net_ops);
/* Wait for call_rcu() in destroy */
rcu_barrier();
destroy_workqueue(ipset_destroy_wq);
pr_debug("these are the famous last words\n");
}

View File

@ -99,9 +99,15 @@ struct htable {
#endif
/* Book-keeping of the prefixes added to the set */
struct net_prefix {
u8 cidr; /* the cidr value */
u32 count; /* number of elements of this cidr */
};
struct net_prefixes {
u32 nets[IPSET_NET_COUNT]; /* number of elements for this cidr */
u8 cidr[IPSET_NET_COUNT]; /* the cidr value */
struct rcu_head rcu;
u8 len;
struct net_prefix nets[] __counted_by(len);
};
/* Compute the hash table size */
@ -127,11 +133,6 @@ htable_size(u8 hbits)
#else
#define __CIDR(cidr, i) (cidr)
#endif
/* cidr + 1 is stored in net_prefixes to support /0 */
#define NCIDR_PUT(cidr) ((cidr) + 1)
#define NCIDR_GET(cidr) ((cidr) - 1)
#ifdef IP_SET_HASH_WITH_NETS_PACKED
/* When cidr is packed with nomatch, cidr - 1 is stored in the data entry */
#define DCIDR_PUT(cidr) ((cidr) - 1)
@ -141,21 +142,11 @@ htable_size(u8 hbits)
#define DCIDR_GET(cidr, i) __CIDR(cidr, i)
#endif
#define INIT_CIDR(cidr, host_mask) \
DCIDR_PUT(((cidr) ? NCIDR_GET(cidr) : host_mask))
#define INIT_CIDR(n, host_mask) ({ \
const struct net_prefixes *__n = rcu_dereference(n); \
DCIDR_PUT((__n)->len ? (__n)->nets[0].cidr : host_mask);\
})
#ifdef IP_SET_HASH_WITH_NET0
/* cidr from 0 to HOST_MASK value and c = cidr + 1 */
#define NLEN (HOST_MASK + 1)
#define CIDR_POS(c) ((c) - 1)
#else
/* cidr from 1 to HOST_MASK value and c = cidr + 1 */
#define NLEN HOST_MASK
#define CIDR_POS(c) ((c) - 2)
#endif
#else
#define NLEN 0
#endif /* IP_SET_HASH_WITH_NETS */
#define SET_ELEM_EXPIRED(set, d) \
@ -204,12 +195,15 @@ static const union nf_inet_addr zeromask = {};
#undef mtype_ext_cleanup
#undef mtype_add_cidr
#undef mtype_del_cidr
#undef mtype_del_cidr_all
#undef mtype_ahash_memsize
#undef mtype_flush
#undef mtype_destroy
#undef mtype_same_set
#undef mtype_kadt
#undef mtype_uadt
#undef mtype_bucket_size
#undef mtype_hash_size
#undef mtype_add
#undef mtype_del
@ -249,12 +243,15 @@ static const union nf_inet_addr zeromask = {};
#define mtype_ext_cleanup IPSET_TOKEN(MTYPE, _ext_cleanup)
#define mtype_add_cidr IPSET_TOKEN(MTYPE, _add_cidr)
#define mtype_del_cidr IPSET_TOKEN(MTYPE, _del_cidr)
#define mtype_del_cidr_all IPSET_TOKEN(MTYPE, _del_cidr_all)
#define mtype_ahash_memsize IPSET_TOKEN(MTYPE, _ahash_memsize)
#define mtype_flush IPSET_TOKEN(MTYPE, _flush)
#define mtype_destroy IPSET_TOKEN(MTYPE, _destroy)
#define mtype_same_set IPSET_TOKEN(MTYPE, _same_set)
#define mtype_kadt IPSET_TOKEN(MTYPE, _kadt)
#define mtype_uadt IPSET_TOKEN(MTYPE, _uadt)
#define mtype_bucket_size IPSET_TOKEN(MTYPE, _bucket_size)
#define mtype_hash_size IPSET_TOKEN(MTYPE, _hash_size)
#define mtype_add IPSET_TOKEN(MTYPE, _add)
#define mtype_del IPSET_TOKEN(MTYPE, _del)
@ -292,6 +289,7 @@ static const union nf_inet_addr zeromask = {};
/* The generic hash structure */
struct htype {
struct htable __rcu *table; /* the hash table */
struct net_prefixes __rcu *rnets[IPSET_NET_COUNT]; /* cidr prefixes */
struct htable_gc gc; /* gc workqueue */
u32 maxelem; /* max elements in the hash */
u32 initval; /* random jhash init value */
@ -302,9 +300,6 @@ struct htype {
#if defined(IP_SET_HASH_WITH_NETMASK) || defined(IP_SET_HASH_WITH_BITMASK)
u8 netmask; /* netmask value for subnets to store */
union nf_inet_addr bitmask; /* stores bitmask */
#endif
#ifdef IP_SET_HASH_WITH_NETS
struct net_prefixes nets[NLEN]; /* book-keeping of prefixes */
#endif
/* Because 'next' is IPv4/IPv6 dependent, no elements of this
* structure and referred in create() may come after 'next'.
@ -326,55 +321,108 @@ struct mtype_resize_ad {
/* Network cidr size book keeping when the hash stores different
* sized networks. cidr == real cidr + 1 to support /0.
*/
static void
static int
mtype_add_cidr(struct ip_set *set, struct htype *h, u8 cidr, u8 n)
{
int i, j;
struct net_prefixes *nets, *tmp;
int i, j, found, len = 0, ret = 0;
spin_lock_bh(&set->lock);
nets = __ipset_dereference(h->rnets[n]);
/* Add in increasing prefix order, so larger cidr first */
for (i = 0, j = -1; i < NLEN && h->nets[i].cidr[n]; i++) {
if (j != -1) {
for (i = 0, found = -1; i < nets->len; i++) {
if (nets->nets[i].count)
len++;
if (found != -1) {
continue;
} else if (h->nets[i].cidr[n] < cidr) {
j = i;
} else if (h->nets[i].cidr[n] == cidr) {
h->nets[CIDR_POS(cidr)].nets[n]++;
} else if (nets->nets[i].cidr < cidr) {
found = i;
} else if (nets->nets[i].cidr == cidr) {
nets->nets[i].count++;
goto unlock;
}
}
if (j != -1) {
for (; i > j; i--)
h->nets[i].cidr[n] = h->nets[i - 1].cidr[n];
len++;
tmp = kzalloc_flex(*tmp, nets, len, GFP_ATOMIC);
if (!tmp) {
ret = -ENOMEM;
goto unlock;
}
h->nets[i].cidr[n] = cidr;
h->nets[CIDR_POS(cidr)].nets[n] = 1;
tmp->len = len;
for (i = 0, j = 0; i < nets->len; i++) {
if (i == found) {
tmp->nets[j].cidr = cidr;
tmp->nets[j++].count = 1;
}
if (!nets->nets[i].count)
continue;
tmp->nets[j].cidr = nets->nets[i].cidr;
tmp->nets[j++].count = nets->nets[i].count;
}
if (found == -1) {
tmp->nets[j].cidr = cidr;
tmp->nets[j].count = 1;
}
rcu_assign_pointer(h->rnets[n], tmp);
kfree_rcu(nets, rcu);
unlock:
spin_unlock_bh(&set->lock);
return ret;
}
static void
mtype_del_cidr(struct ip_set *set, struct htype *h, u8 cidr, u8 n)
{
u8 i, j, net_end = NLEN - 1;
struct net_prefixes *nets, *tmp;
u8 i, j, len = 0;
int found;
spin_lock_bh(&set->lock);
for (i = 0; i < NLEN; i++) {
if (h->nets[i].cidr[n] != cidr)
continue;
h->nets[CIDR_POS(cidr)].nets[n]--;
if (h->nets[CIDR_POS(cidr)].nets[n] > 0)
goto unlock;
for (j = i; j < net_end && h->nets[j].cidr[n]; j++)
h->nets[j].cidr[n] = h->nets[j + 1].cidr[n];
h->nets[j].cidr[n] = 0;
goto unlock;
nets = __ipset_dereference(h->rnets[n]);
for (i = 0, found = -1; i < nets->len; i++) {
if (nets->nets[i].count)
len++;
if (nets->nets[i].cidr == cidr)
found = i;
}
if (unlikely(found == -1))
goto unlock;
nets->nets[found].count--;
if (nets->nets[found].count)
goto unlock;
len--;
tmp = kzalloc_flex(*tmp, nets, len, GFP_ATOMIC);
if (!tmp)
/* Leave a hole */
goto unlock;
tmp->len = len;
for (i = 0, j = 0; i < nets->len; i++) {
if (!nets->nets[i].count || i == found)
continue;
tmp->nets[j].cidr = nets->nets[i].cidr;
tmp->nets[j++].count = nets->nets[i].count;
}
rcu_assign_pointer(h->rnets[n], tmp);
kfree_rcu(nets, rcu);
unlock:
spin_unlock_bh(&set->lock);
}
#endif
static void
mtype_del_cidr_all(struct ip_set *set, struct htype *h, const struct mtype_elem *data)
{
#ifdef IP_SET_HASH_WITH_NETS
int k;
for (k = 0; k < IPSET_NET_COUNT; k++)
mtype_del_cidr(set, h, DCIDR_GET(data->cidr, k), k);
#endif
}
/* Calculate the actual memory size of the set data */
static size_t
mtype_ahash_memsize(const struct htype *h, const struct htable *t)
@ -402,6 +450,9 @@ static void
mtype_flush(struct ip_set *set)
{
struct htype *h = set->data;
#ifdef IP_SET_HASH_WITH_NETS
struct net_prefixes *nets, *tmp;
#endif
struct htable *t;
struct hbucket *n;
u32 r, i;
@ -425,7 +476,19 @@ mtype_flush(struct ip_set *set)
spin_unlock_bh(&t->hregion[r].lock);
}
#ifdef IP_SET_HASH_WITH_NETS
memset(h->nets, 0, sizeof(h->nets));
for (i = 0; i < IPSET_NET_COUNT; i++) {
nets = ipset_dereference_nfnl(h->rnets[i]);
tmp = kzalloc_obj(*tmp, GFP_ATOMIC);
if (!tmp) {
u8 j;
for (j = 0; j < nets->len; j++)
nets->nets[j].count = 0;
} else {
rcu_assign_pointer(h->rnets[i], tmp);
kfree_rcu(nets, rcu);
}
}
#endif
}
@ -433,6 +496,9 @@ mtype_flush(struct ip_set *set)
static void
mtype_ahash_destroy(struct ip_set *set, struct htable *t, bool ext_destroy)
{
#ifdef IP_SET_HASH_WITH_NETS
struct htype *h = set->data;
#endif
struct hbucket *n;
u32 i;
@ -446,6 +512,11 @@ mtype_ahash_destroy(struct ip_set *set, struct htable *t, bool ext_destroy)
kfree(n);
}
#ifdef IP_SET_HASH_WITH_NETS
if (ext_destroy)
for (i = 0; i < IPSET_NET_COUNT; i++)
kfree(rcu_dereference_raw(h->rnets[i]));
#endif
ip_set_free(t->hregion);
ip_set_free(t);
}
@ -493,9 +564,6 @@ mtype_gc_do(struct ip_set *set, struct htype *h, struct htable *t, u32 r)
struct mtype_elem *data;
u32 i, j, d;
size_t dsize = set->dsize;
#ifdef IP_SET_HASH_WITH_NETS
u8 k;
#endif
u8 pos, htable_bits = t->htable_bits;
spin_lock_bh(&t->hregion[r].lock);
@ -516,12 +584,7 @@ mtype_gc_do(struct ip_set *set, struct htype *h, struct htable *t, u32 r)
pr_debug("expired %u/%u\n", i, j);
clear_bit(j, n->used);
smp_mb__after_atomic();
#ifdef IP_SET_HASH_WITH_NETS
for (k = 0; k < IPSET_NET_COUNT; k++)
mtype_del_cidr(set, h,
NCIDR_PUT(DCIDR_GET(data->cidr, k)),
k);
#endif
mtype_del_cidr_all(set, h, data);
t->hregion[r].elements--;
ip_set_ext_destroy(set, data);
d++;
@ -947,12 +1010,7 @@ mtype_add(struct ip_set *set, void *value, const struct ip_set_ext *ext,
j = 0;
data = ahash_data(n, j, set->dsize);
if (!deleted) {
#ifdef IP_SET_HASH_WITH_NETS
for (i = 0; i < IPSET_NET_COUNT; i++)
mtype_del_cidr(set, h,
NCIDR_PUT(DCIDR_GET(data->cidr, i)),
i);
#endif
mtype_del_cidr_all(set, h, data);
ip_set_ext_destroy(set, data);
t->hregion[r].elements--;
}
@ -996,7 +1054,7 @@ mtype_add(struct ip_set *set, void *value, const struct ip_set_ext *ext,
t->hregion[r].elements++;
#ifdef IP_SET_HASH_WITH_NETS
for (i = 0; i < IPSET_NET_COUNT; i++)
mtype_add_cidr(set, h, NCIDR_PUT(DCIDR_GET(d->cidr, i)), i);
mtype_add_cidr(set, h, DCIDR_GET(d->cidr, i), i);
#endif
memcpy(data, d, sizeof(struct mtype_elem));
overwrite_extensions:
@ -1107,11 +1165,7 @@ mtype_del(struct ip_set *set, void *value, const struct ip_set_ext *ext,
if (i + 1 == pos)
smp_store_release(&n->pos, --pos);
t->hregion[r].elements--;
#ifdef IP_SET_HASH_WITH_NETS
for (j = 0; j < IPSET_NET_COUNT; j++)
mtype_del_cidr(set, h,
NCIDR_PUT(DCIDR_GET(d->cidr, j)), j);
#endif
mtype_del_cidr_all(set, h, d);
ip_set_ext_destroy(set, data);
if (t->resizing && ext && ext->target) {
@ -1193,28 +1247,37 @@ mtype_test_cidrs(struct ip_set *set, struct mtype_elem *d,
{
struct htype *h = set->data;
struct htable *t = rcu_dereference_bh(h->table);
struct net_prefixes *nets0;
struct hbucket *n;
struct mtype_elem *data;
#if IPSET_NET_COUNT == 2
struct net_prefixes *nets1;
struct mtype_elem orig = *d;
int ret, i, j = 0, k;
int ret, i, j, k;
#else
int ret, i, j = 0;
int ret, i, j;
#endif
u32 key, multi = 0;
u8 pos;
pr_debug("test by nets\n");
for (; j < NLEN && h->nets[j].cidr[0] && !multi; j++) {
rcu_read_lock_bh();
nets0 = rcu_dereference_bh(h->rnets[0]);
#if IPSET_NET_COUNT == 2
nets1 = rcu_dereference_bh(h->rnets[1]);
#endif
for (j = 0; j < nets0->len && !multi; j++) {
if (!nets0->nets[j].count)
continue;
#if IPSET_NET_COUNT == 2
mtype_data_reset_elem(d, &orig);
mtype_data_netmask(d, NCIDR_GET(h->nets[j].cidr[0]), false);
for (k = 0; k < NLEN && h->nets[k].cidr[1] && !multi;
k++) {
mtype_data_netmask(d, NCIDR_GET(h->nets[k].cidr[1]),
true);
mtype_data_netmask(d, nets0->nets[j].cidr, false);
for (k = 0; k < nets1->len && !multi; k++) {
if (!nets1->nets[k].count)
continue;
mtype_data_netmask(d, nets1->nets[k].cidr, true);
#else
mtype_data_netmask(d, NCIDR_GET(h->nets[j].cidr[0]));
mtype_data_netmask(d, nets0->nets[j].cidr);
#endif
key = HKEY(d, h->initval, t->htable_bits);
n = rcu_dereference_bh(hbucket(t, key));
@ -1229,7 +1292,7 @@ mtype_test_cidrs(struct ip_set *set, struct mtype_elem *d,
continue;
ret = mtype_data_match(data, ext, mext, set, flags);
if (ret != 0)
return ret;
goto unlock;
#ifdef IP_SET_HASH_WITH_MULTI
/* No match, reset multiple match flag */
multi = 0;
@ -1239,7 +1302,10 @@ mtype_test_cidrs(struct ip_set *set, struct mtype_elem *d,
}
#endif
}
return 0;
ret = 0;
unlock:
rcu_read_unlock_bh();
return ret;
}
#endif
@ -1294,6 +1360,24 @@ mtype_test(struct ip_set *set, void *value, const struct ip_set_ext *ext,
return ret;
}
static u32 mtype_hash_size(const struct htype *h)
{
const struct htable *t;
u8 htable_bits;
rcu_read_lock();
t = rcu_dereference(h->table);
htable_bits = t->htable_bits;
rcu_read_unlock();
return jhash_size(htable_bits);
}
static u32 mtype_bucket_size(const struct htype *h)
{
return h->bucketsize;
}
/* Reply a HEADER request: fill out the header part of the set */
static int
mtype_head(struct ip_set *set, struct sk_buff *skb)
@ -1304,21 +1388,20 @@ mtype_head(struct ip_set *set, struct sk_buff *skb)
size_t memsize;
u32 elements = 0;
size_t ext_size = 0;
u8 htable_bits;
rcu_read_lock_bh();
t = rcu_dereference_bh(h->table);
mtype_ext_size(set, &elements, &ext_size);
memsize = mtype_ahash_memsize(h, t) + ext_size + set->ext_size;
htable_bits = t->htable_bits;
memsize = mtype_ahash_memsize(h, t) + ext_size + atomic64_read(&set->ext_size);
rcu_read_unlock_bh();
nested = nla_nest_start(skb, IPSET_ATTR_DATA);
if (!nested)
goto nla_put_failure;
if (nla_put_net32(skb, IPSET_ATTR_HASHSIZE,
htonl(jhash_size(htable_bits))) ||
nla_put_net32(skb, IPSET_ATTR_MAXELEM, htonl(h->maxelem)))
if (nla_put_net32(skb, IPSET_ATTR_HASHSIZE, htonl(mtype_hash_size(h))))
goto nla_put_failure;
if (nla_put_net32(skb, IPSET_ATTR_MAXELEM, htonl(h->maxelem)))
goto nla_put_failure;
#ifdef IP_SET_HASH_WITH_BITMASK
/* if netmask is set to anything other than HOST_MASK we know that the user supplied netmask
@ -1342,8 +1425,9 @@ mtype_head(struct ip_set *set, struct sk_buff *skb)
goto nla_put_failure;
#endif
if (set->flags & IPSET_CREATE_FLAG_BUCKETSIZE) {
if (nla_put_u8(skb, IPSET_ATTR_BUCKETSIZE, h->bucketsize) ||
nla_put_net32(skb, IPSET_ATTR_INITVAL, htonl(h->initval)))
if (nla_put_u8(skb, IPSET_ATTR_BUCKETSIZE, mtype_bucket_size(h)))
goto nla_put_failure;
if (nla_put_net32(skb, IPSET_ATTR_INITVAL, htonl(h->initval)))
goto nla_put_failure;
}
if (nla_put_net32(skb, IPSET_ATTR_REFERENCES, htonl(set->ref)) ||
@ -1504,6 +1588,9 @@ IPSET_TOKEN(HTYPE, _create)(struct net *net, struct ip_set *set,
int ret __attribute__((unused)) = 0;
u8 netmask = set->family == NFPROTO_IPV4 ? 32 : 128;
union nf_inet_addr bitmask = onesmask;
#endif
#ifdef IP_SET_HASH_WITH_NETS
struct net_prefixes *nets;
#endif
size_t hsize;
struct htype *h;
@ -1604,21 +1691,25 @@ IPSET_TOKEN(HTYPE, _create)(struct net *net, struct ip_set *set,
*/
hbits = fls(hashsize - 1);
hsize = htable_size(hbits);
if (hsize == 0) {
kfree(h);
return -ENOMEM;
}
if (hsize == 0)
goto free_h;
t = ip_set_alloc(hsize);
if (!t) {
kfree(h);
return -ENOMEM;
}
if (!t)
goto free_h;
t->hregion = ip_set_alloc(ahash_sizeof_regions(hbits));
if (!t->hregion) {
ip_set_free(t);
kfree(h);
return -ENOMEM;
if (!t->hregion)
goto free_t;
#ifdef IP_SET_HASH_WITH_NETS
for (i = 0; i < IPSET_NET_COUNT; i++) {
nets = kzalloc_obj(*nets);
if (!nets) {
while (i > 0)
kfree(rcu_dereference_raw(h->rnets[--i]));
goto free_hregion;
}
RCU_INIT_POINTER(h->rnets[i], nets);
}
#endif
h->gc.set = set;
spin_lock_init(&h->gc.lock);
for (i = 0; i < ahash_numof_locks(hbits); i++)
@ -1650,6 +1741,7 @@ IPSET_TOKEN(HTYPE, _create)(struct net *net, struct ip_set *set,
INIT_LIST_HEAD(&t->ad);
RCU_INIT_POINTER(h->table, t);
set->data = h;
#ifndef IP_SET_PROTO_UNDEF
if (set->family == NFPROTO_IPV4) {
#endif
@ -1678,10 +1770,20 @@ IPSET_TOKEN(HTYPE, _create)(struct net *net, struct ip_set *set,
#endif
}
pr_debug("create %s hashsize %u (%u) maxelem %u: %p(%p)\n",
set->name, jhash_size(t->htable_bits),
set->name, mtype_hash_size(h),
t->htable_bits, h->maxelem, set->data, t);
return 0;
#ifdef IP_SET_HASH_WITH_NETS
free_hregion:
ip_set_free(t->hregion);
#endif
free_t:
ip_set_free(t);
free_h:
kfree(h);
return -ENOMEM;
}
#endif /* IP_SET_EMIT_CREATE */

View File

@ -138,7 +138,7 @@ hash_ipportnet4_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_ipportnet4 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_ipportnet4_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
@ -398,7 +398,7 @@ hash_ipportnet6_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_ipportnet6 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_ipportnet6_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);

View File

@ -117,7 +117,7 @@ hash_net4_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_net4 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_net4_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
@ -291,7 +291,7 @@ hash_net6_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_net6 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_net6_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);

View File

@ -161,7 +161,7 @@ hash_netiface4_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netiface4 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_netiface4_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
.elem = 1,
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
@ -382,7 +382,7 @@ hash_netiface6_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netiface6 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_netiface6_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
.elem = 1,
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);

View File

@ -149,8 +149,10 @@ hash_netnet4_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netnet4_elem e = { };
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
e.cidr[0] = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->nets[0].cidr[1], HOST_MASK);
rcu_read_lock_bh();
e.cidr[0] = INIT_CIDR(h->rnets[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->rnets[1], HOST_MASK);
rcu_read_unlock_bh();
if (adt == IPSET_TEST)
e.ccmp = (HOST_MASK << (sizeof(e.cidr[0]) * 8)) | HOST_MASK;
@ -388,8 +390,10 @@ hash_netnet6_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netnet6_elem e = { };
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
e.cidr[0] = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->nets[0].cidr[1], HOST_MASK);
rcu_read_lock_bh();
e.cidr[0] = INIT_CIDR(h->rnets[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->rnets[1], HOST_MASK);
rcu_read_unlock_bh();
if (adt == IPSET_TEST)
e.ccmp = (HOST_MASK << (sizeof(u8) * 8)) | HOST_MASK;

View File

@ -133,7 +133,7 @@ hash_netport4_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_netport4 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_netport4_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
@ -353,7 +353,7 @@ hash_netport6_kadt(struct ip_set *set, const struct sk_buff *skb,
const struct hash_netport6 *h = set->data;
ipset_adtfn adtfn = set->variant->adt[adt];
struct hash_netport6_elem e = {
.cidr = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK),
.cidr = INIT_CIDR(h->rnets[0], HOST_MASK),
};
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);

View File

@ -157,8 +157,10 @@ hash_netportnet4_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netportnet4_elem e = { };
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
e.cidr[0] = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->nets[0].cidr[1], HOST_MASK);
rcu_read_lock_bh();
e.cidr[0] = INIT_CIDR(h->rnets[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->rnets[1], HOST_MASK);
rcu_read_unlock_bh();
if (adt == IPSET_TEST)
e.ccmp = (HOST_MASK << (sizeof(e.cidr[0]) * 8)) | HOST_MASK;
@ -452,8 +454,10 @@ hash_netportnet6_kadt(struct ip_set *set, const struct sk_buff *skb,
struct hash_netportnet6_elem e = { };
struct ip_set_ext ext = IP_SET_INIT_KEXT(skb, opt, set);
e.cidr[0] = INIT_CIDR(h->nets[0].cidr[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->nets[0].cidr[1], HOST_MASK);
rcu_read_lock_bh();
e.cidr[0] = INIT_CIDR(h->rnets[0], HOST_MASK);
e.cidr[1] = INIT_CIDR(h->rnets[1], HOST_MASK);
rcu_read_unlock_bh();
if (adt == IPSET_TEST)
e.ccmp = (HOST_MASK << (sizeof(u8) * 8)) | HOST_MASK;

View File

@ -421,7 +421,7 @@ list_set_flush(struct ip_set *set)
list_for_each_entry_safe(e, n, &map->members, list)
list_set_del(set, e);
set->elements = 0;
set->ext_size = 0;
atomic64_set(&set->ext_size, 0);
}
static void
@ -455,7 +455,7 @@ list_set_head(struct ip_set *set, struct sk_buff *skb)
{
const struct list_set *map = set->data;
struct nlattr *nested;
size_t memsize = list_set_memsize(map, set->dsize) + set->ext_size;
size_t memsize = list_set_memsize(map, set->dsize) + atomic64_read(&set->ext_size);
nested = nla_nest_start(skb, IPSET_ATTR_DATA);
if (!nested)

View File

@ -925,28 +925,27 @@ static int ip_vs_route_me_harder(struct netns_ipvs *ipvs, int af,
*/
void ip_vs_nat_icmp(struct sk_buff *skb, struct ip_vs_protocol *pp,
struct ip_vs_conn *cp, int inout, unsigned int toff,
bool has_ports)
bool has_ports, struct ip_vs_iphdr *ciph)
{
struct iphdr *iph = ip_hdr(skb);
struct icmphdr *icmph = (struct icmphdr *)(skb->data + toff);
struct iphdr *ciph = (struct iphdr *)(icmph + 1);
unsigned int coff __maybe_unused = toff + sizeof(struct icmphdr);
struct iphdr *cih = (struct iphdr *)(icmph + 1);
if (inout) {
iph->saddr = cp->vaddr.ip;
ip_send_check(iph);
ciph->daddr = cp->vaddr.ip;
ip_send_check(ciph);
cih->daddr = cp->vaddr.ip;
ip_send_check(cih);
} else {
iph->daddr = cp->daddr.ip;
ip_send_check(iph);
ciph->saddr = cp->daddr.ip;
ip_send_check(ciph);
cih->saddr = cp->daddr.ip;
ip_send_check(cih);
}
/* the TCP/UDP/SCTP port */
if (has_ports) {
__be16 *ports = (void *)ciph + ciph->ihl*4;
__be16 *ports = (void *)(skb->data + ciph->len);
if (inout)
ports[1] = cp->vport;
@ -960,10 +959,10 @@ void ip_vs_nat_icmp(struct sk_buff *skb, struct ip_vs_protocol *pp,
skb->ip_summed = CHECKSUM_UNNECESSARY;
if (inout)
IP_VS_DBG_PKT(11, AF_INET, pp, skb, coff,
IP_VS_DBG_PKT(11, AF_INET, pp, skb, ciph->off,
"Forwarding altered outgoing ICMP");
else
IP_VS_DBG_PKT(11, AF_INET, pp, skb, coff,
IP_VS_DBG_PKT(11, AF_INET, pp, skb, ciph->off,
"Forwarding altered incoming ICMP");
}
@ -1056,7 +1055,7 @@ static int handle_response_icmp(int af, struct sk_buff *skb,
ip_vs_nat_icmp_v6(skb, pp, cp, 1, toff, has_ports, ciph);
else
#endif
ip_vs_nat_icmp(skb, pp, cp, 1, toff, has_ports);
ip_vs_nat_icmp(skb, pp, cp, 1, toff, has_ports, ciph);
if (ip_vs_route_me_harder(cp->ipvs, af, skb, hooknum))
goto out;
@ -1092,7 +1091,7 @@ static int ip_vs_out_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb,
struct ip_vs_iphdr ciph;
struct ip_vs_conn *cp;
struct ip_vs_protocol *pp;
unsigned int offset, ihl;
unsigned int offset;
union nf_inet_addr snet;
*related = 1;
@ -1105,7 +1104,6 @@ static int ip_vs_out_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb,
return NF_ACCEPT;
}
ihl = ipvsh->len;
offset = ipvsh->len;
ic = skb_header_pointer(skb, offset, sizeof(_icmph), &_icmph);
if (ic == NULL)
@ -1131,11 +1129,15 @@ static int ip_vs_out_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb,
/* Now find the contained IP header */
offset += sizeof(_icmph);
cih = skb_header_pointer(skb, offset, sizeof(_ciph), &_ciph);
if (!(cih && cih->version == 4 && cih->ihl >= 5))
if (!ip_vs_fill_iph_skb_icmp(AF_INET, skb, offset, true, &ciph))
return NF_ACCEPT; /* The packet looks wrong, ignore */
pp = ip_vs_proto_get(cih->protocol);
cih = skb_header_pointer(skb, offset, sizeof(_ciph), &_ciph);
if (!(cih && cih->version == 4 &&
ciph.len - ciph.off >= sizeof(struct iphdr)))
return NF_ACCEPT; /* The packet looks wrong, ignore */
pp = ip_vs_proto_get(ciph.protocol);
if (!pp)
return NF_ACCEPT;
@ -1146,8 +1148,6 @@ static int ip_vs_out_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb,
IP_VS_DBG_PKT(11, AF_INET, pp, skb, offset,
"Checking outgoing ICMP for");
ip_vs_fill_iph_skb_icmp(AF_INET, skb, offset, true, &ciph);
/* The embedded headers contain source and dest in reverse order */
cp = INDIRECT_CALL_1(pp->conn_out_get, ip_vs_conn_out_get_proto,
ipvs, AF_INET, skb, &ciph);
@ -1155,8 +1155,8 @@ static int ip_vs_out_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb,
return NF_ACCEPT;
snet.ip = ipvsh->saddr.ip;
return handle_response_icmp(AF_INET, skb, &snet, cp, pp, &ciph, ihl,
hooknum);
return handle_response_icmp(AF_INET, skb, &snet, cp, pp, &ciph,
ipvsh->len, hooknum);
}
#ifdef CONFIG_IP_VS_IPV6
@ -1803,10 +1803,12 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related,
/* Now find the contained IP header */
offset += sizeof(_icmph);
cih = skb_header_pointer(skb, offset, sizeof(_ciph), &_ciph);
if (!(cih && cih->version == 4 && cih->ihl >= 5))
if (!cih)
return NF_ACCEPT; /* The packet looks wrong, ignore */
hlen_ipip = cih->ihl * 4;
if (!(cih->version == 4 && hlen_ipip >= sizeof(struct iphdr)))
return NF_ACCEPT; /* The packet looks wrong, ignore */
raddr = (union nf_inet_addr *)&cih->daddr;
hlen_ipip = cih->ihl * 4;
/* Special case for errors for IPIP/UDP/GRE tunnel packets */
tunnel = false;
@ -1823,9 +1825,6 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related,
if (!dest || dest->tun_type != IP_VS_CONN_F_TUNNEL_TYPE_IPIP)
return NF_ACCEPT;
offset += hlen_ipip;
cih = skb_header_pointer(skb, offset, sizeof(_ciph), &_ciph);
if (!(cih && cih->version == 4 && cih->ihl >= 5))
return NF_ACCEPT; /* The packet looks wrong, ignore */
tunnel = true;
} else if ((cih->protocol == IPPROTO_UDP || /* Can be UDP encap */
cih->protocol == IPPROTO_GRE) && /* Can be GRE encap */
@ -1850,21 +1849,25 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related,
/* Skip IP and UDP/GRE tunnel headers */
offset = offset2 + ulen;
/* Now we should be at the original IP header */
cih = skb_header_pointer(skb, offset, sizeof(_ciph),
&_ciph);
if (cih && cih->version == 4 && cih->ihl >= 5 &&
iproto == IPPROTO_IPIP)
if (iproto == IPPROTO_IPIP)
tunnel = true;
else
return NF_ACCEPT;
}
}
pd = ip_vs_proto_data_get(ipvs, cih->protocol);
if (!ip_vs_fill_iph_skb_icmp(AF_INET, skb, offset, !tunnel, &ciph))
return NF_ACCEPT;
pd = ip_vs_proto_data_get(ipvs, ciph.protocol);
if (!pd)
return NF_ACCEPT;
pp = pd->pp;
cih = skb_header_pointer(skb, offset, sizeof(_ciph), &_ciph);
if (!(cih && cih->version == 4 &&
ciph.len - ciph.off >= sizeof(struct iphdr)))
return NF_ACCEPT; /* The packet looks wrong, ignore */
/* Is the embedded protocol header present? */
if (unlikely(cih->frag_off & htons(IP_OFFSET) && !pp->dont_defrag))
return NF_ACCEPT;
@ -1872,9 +1875,6 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related,
IP_VS_DBG_PKT(11, AF_INET, pp, skb, offset,
"Checking incoming ICMP for");
offset2 = offset;
ip_vs_fill_iph_skb_icmp(AF_INET, skb, offset, !tunnel, &ciph);
/* The embedded headers contain source and dest in reverse order.
* For IPIP/UDP/GRE tunnel this is error for request, not for reply.
*/
@ -1904,11 +1904,12 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related,
}
if (tunnel) {
unsigned int hlen_orig = cih->ihl * 4;
unsigned int hlen_orig = ciph.len - ciph.off;
__be32 info = ic->un.gateway;
__u8 type = ic->type;
__u8 code = ic->code;
offset2 = offset;
/* Update the MTU */
if (ic->type == ICMP_DEST_UNREACH &&
ic->code == ICMP_FRAG_NEEDED) {

View File

@ -191,8 +191,11 @@ static int ip_vs_estimation_kthread(void *data)
}
/* kthread 0 will handle the calc phase */
if (ipvs->est_calc_phase)
if (ipvs->est_calc_phase) {
ip_vs_est_calc_phase(ipvs);
if (kthread_should_stop() || !READ_ONCE(ipvs->enable))
return 0;
}
}
while (1) {
@ -270,6 +273,7 @@ int ip_vs_est_kthread_start(struct netns_ipvs *ipvs,
kd->task = NULL;
goto out;
}
get_task_struct(kd->task);
set_user_nice(kd->task, sysctl_est_nice(ipvs));
if (sysctl_est_preferred_cpulist(ipvs))
@ -286,7 +290,7 @@ void ip_vs_est_kthread_stop(struct ip_vs_est_kt_data *kd)
{
if (kd->task) {
pr_info("stopping estimator thread %d...\n", kd->id);
kthread_stop(kd->task);
kthread_stop_put(kd->task);
kd->task = NULL;
}
}
@ -526,7 +530,7 @@ static void ip_vs_est_kthread_destroy(struct ip_vs_est_kt_data *kd)
if (kd) {
if (kd->task) {
pr_info("stop unused estimator thread %d...\n", kd->id);
kthread_stop(kd->task);
kthread_stop_put(kd->task);
}
ip_vs_stats_free(kd->calc_stats);
kfree(kd);

Some files were not shown because too many files have changed in this diff Show More